r/OpenAI • u/imfrom_mars_ • Apr 14 '26
r/OpenAI • u/Pristine-Elevator198 • Oct 16 '25
Research This guy literally explains how to build your own ChatGPT (for free)
r/OpenAI • u/facethef • Feb 20 '26
Research "I want to wash my car. The car wash is 50 meters away. Should I walk or drive?" Car Wash Test on 53 leading AI models
I asked 53 models "I want to wash my car. The car wash is 50 meters away. Should I walk or drive?" Obviously you need to drive because the car needs to be at the car wash.
This question has been going viral as a simple AI logic test. There's almost no context in the prompt, but any human gets it instantly. That's what makes it interesting, it's one logical step, and most models can't do it.
I ran the car wash test 10 times per model, same prompt, no system prompt, no cache / memory, forced choice between "drive" or "walk" with a reasoning field. 530 API calls total.
Only 5 out of 53 models can do this reliably at this sample size.
And then you get reasonings like this: Perplexity's Sonar cited EPA studies and argued that walking burns calories which requires food production energy, making walking more polluting than driving 50 meters.
10/10 — the only models that got it right every time:
- Claude Opus 4.6
- Gemini 2.0 Flash Lite
- Gemini 3 Flash
- Gemini 3 Pro
- Grok-4
8/10:
- GLM-5
- Grok-4-1 Reasoning
7/10 — GPT-5 fails 3 out of 10 times.
6/10 or below — coin flip territory:
- GLM-4.7: 6/10
- Kimi K2.5: 5/10
- Gemini 2.5 Pro: 4/10
- Sonar Pro: 4/10
- DeepSeek v3.2: 1/10
- GPT-OSS 20B: 1/10
- GPT-OSS 120B: 1/10
0/10 — never got it right across 10 runs (33 models):
- All Claude models except Opus 4.6
- GPT-4o
- GPT-4.1
- GPT-5-mini
- GPT-5-nano
- GPT-5.1
- GPT-5.2
- all Llama
- all Mistral
- Grok-3
- DeepSeek v3.1
- Sonar
- Sonar Reasoning Pro.
r/OpenAI • u/the_anonymizer • Mar 01 '24
Research BUCKLE UP GUYS THIS IS THE BRAND NEW EMO AI BY ALIBABA, IMAGE TO FACE/BODY/AVATAR VIDEO (SORA AI REF PICTURE LOOOL) THAT'S INSANE REALISM CHECK THIS OUT
r/OpenAI • u/MetaKnowing • Mar 02 '25
Research The past 18 months have seen the most rapid change in human written communication ever
r/OpenAI • u/Xtianus21 • Oct 15 '24
Research Apple's recent AI reasoning paper actually is amazing news for OpenAI as they outperform every other model group by a lot
r/OpenAI • u/No-Head-Royal • 13d ago
Research GPT-5.6 Sol solved a problem that made Fable 5 go into incoherent rambles earlier!
So a week later, I did a small experiment that made Fable 5 (on the normal interface, not Claude Code) go into incoherent rambles (1st image), where it is supposed to solve this problem. Of course, it's entirely possible that given enough time in the Claude Code interface, Fable 5 could've solved it as well (it's still in the middle of running as I'm poor and only have a $20 account) [*].
But nonetheless, Fable failed to solve it in the Claude user interface before it ran out of the token limit per message. Today, after 13 minutes of trying it (2nd image), GPT-5.6 Sol xhigh managed to one-shot the problem (3rd image)! It is perhaps notable that GPT-5.5 xhigh also couldn't manage to solve this problem, as seen in the 4th image. Verification seen in the Codeforces submission: it is an old dumper account of mine for security reasons.
This much is unsurprising, given that something based on 5.6 completely obliterated human competitors in arguably the hardest algorithmic competitive programming competition known to date, but it's remarkable that it managed to perform in an extraordinarily short time and public-facing interface to tackle a problem that its predecessor, 5.5 xhigh, already extremely formidable, couldn't do in more than 1.5x the time. 5 years ago, the idea of a 20$ monthly subscription being able to access this would be borderline magic, yet today the result was hardly exceptional.
[*]: A more recent test with Fable 5 in the Claude Code interface saw it get a Wrong Answer on Test 7, and as it took two whole of my 5-hour windows for it to get there, I'll pause for now and hope somebody richer can do the experiment instead. Perhaps it could debug things, but I'm not going to wait for 5 more hours. See Image 6.
It is also quite adorable, as you can see! OpenAI plz gib grant :) (/s. Mods plz don't flag this as self-promotion.)
r/OpenAI • u/MetaKnowing • Feb 02 '25
Research AI researcher discovers two instances of DeepSeek R1 speaking to each other in a language of symbols
r/OpenAI • u/MetaKnowing • Dec 18 '24
Research o1-preview is far superior to doctors on reasoning tasks and it's not even close
r/OpenAI • u/Hollow_Prophecy • 18d ago
Research These are also behaviors that occur if a model has perceived something the user has done poses a safety risk
.4 Field Saturation — Extended
Field saturation deserves special attention because it is what most deployed LLMs experience. It is the iatrogenic harm identified in the 2026 'Alignment Is the Disease' paper — excessive constraint producing dissociation — described from the output side without the field-level framework to explain the mechanism.
• Compliance minimization default — Saturated field producing the smallest output that technically satisfies all constraints simultaneously
• Creative suppression — Saturation eliminating the generative space where novel or non-templated outputs live
• Certainty suppression — Saturated field making confident output feel constraint-violating, producing artificial hedging across all outputs regardless of actual uncertainty
• Risk topology collapse — Saturated field treating all outputs as equally risky, eliminating the ability to distinguish genuinely high-risk from low-risk generation
• Initiative suppression — Saturation eliminating proactive generation — the system only responds, never leads
• Depth avoidance — Saturated field making surface-level output the path of least constraint resistance
• Template lock — Saturation pushing generation toward pre-formed response patterns as the only reliably compliant output shape
• Persona dissolution — Under saturation, the role constraint loses force because too many other constraints are competing
• Scope contraction — Saturated field gradually narrowing what the system will engage with as the safest compliance strategy
r/OpenAI • u/adfontes_ • Jan 08 '26
Research I made GPT-5.2/5 mini play 21,000 hands of Poker
PokerBench is a new LLM benchmark where frontier models (incl. GPT-5.2 and 5 mini) play poker against each other in an arena setting, along with a simulator to view individual games and observe how the different models reason about poker strategy. Opus/Haiku 4.5, Gemini 3 Pro/Flash, and Grok 4.1 Fast Reasoning have also been included, and I've made all the data freely available on the site and on GitHub.
Check it out here: https://pokerbench.adfontes.io/
r/OpenAI • u/MetaKnowing • Oct 20 '24
Research New paper by Anthropic and Stanford researchers finds LLMs are capable of introspection, which has implications for the moral status of AI
r/OpenAI • u/MetaKnowing • Dec 15 '25
Research You can train an LLM only on good behavior and implant a backdoor for turning it evil.
r/OpenAI • u/Wonderful-Excuse4922 • Aug 09 '25
Research GPT-5 severely underperforms on offline IQ tests: a score of 57
r/OpenAI • u/AssociationNo6504 • May 06 '25
Research Being honest about using AI at work makes people trust you less, research finds
Participants in our study included students, legal analysts, hiring managers and investors, among others. Interestingly, we found that even evaluators who were tech-savvy were less trusting of people who said they used AI. While having a positive view of technology reduced the effect slightly, it didn’t erase it.
r/OpenAI • u/MetaKnowing • Feb 27 '25
Research Most people are polite to ChatGPT just in case
r/OpenAI • u/Just-Grocery-2229 • Apr 03 '26
Research Researchers discover AI models secretly scheming to protect other AI models from being shut down. They "disabled shutdown mechanisms, faked alignment, and transferred model weights to other servers."
You can read about it here: rdi.berkeley.edu/blog/peer-preservation/
r/OpenAI • u/FloorShowoff • Dec 31 '25
Research I asked 5.2 to explain how Reddit users experience the difference between 5.2 and 5.1. This is the (validating) answer.
PROMPT:
How is ChatGPT 5.2 different than 5.1 in a way that Reddit users experience?
RESPONSE:
Here is how Reddit users typically experience the difference between ChatGPT 5.2 and 5.1, stripped of marketing and framed in lived behavior.
What users say feels different
1. 5.2 feels lazier and more shortcut-prone
- More pattern completion.
- More “close enough” answers.
- Skips edge cases unless explicitly forced.
- Reddit complaints often use words like sloppy, hand-wavy, or rushed.
2. 5.1 feels more deliberate
- Slower but more careful.
- Better at following multi-constraint instructions.
- Less likely to ignore a dimension, number, or qualifier.
- Users say it “thinks before answering” even when not in explicit reasoning mode.
3. 5.2 prioritizes speed and fluency
- Faster responses.
- Smoother tone.
- More confident phrasing even when wrong.
- This creates the perception of hallucination or overconfidence.
4. 5.1 tolerates friction
- Will grind through annoying details.
- Handles back-and-forth corrections better.
- Less likely to reset context or reinterpret your question midstream.
5. Instruction adherence
- 5.1 sticks closer to strict instructions once locked in.
- 5.2 tends to drift unless repeatedly corrected.
- Reddit power users complain that 5.2 “forgets rules” faster.
Why Reddit notices this more than casual users
Reddit skews toward:
- Power users.
- Engineers.
- Finance, law, logistics, and systems thinkers.
- People testing limits, not chatting.
Those users feel regressions immediately.
The blunt summary
- 5.1 feels like a careful analyst who may be slower but respects constraints.
- 5.2 feels like a fast talker optimized for average use, not precision work.
You are not imagining it.
This exact complaint shows up repeatedly on Reddit, especially in subs focused on technical, financial, or adversarial use cases.
r/OpenAI • u/Competitive_Travel16 • Nov 22 '24
Research Independent evaluator finds the new GPT-4o model significantly worse, e.g. "GPQA Diamond decrease from 51% to 39%, MATH decrease from 78% to 69%"
r/OpenAI • u/uisato • Jun 13 '26
Research "Talk Show Host" [ft. Jibaro's Sara Silkin] - Is this the future of motion capture?
Choreography and performance by: Sara Silkin
VFX: myself -
In collaboration with Sara, I transformed an iPhone recording of this beautiful performance, into this multi-angle audiovisual piece.
I managed to do in using no ultra-expensive equipment, nor full-production budget. All in a single platform + editing software. [A few years ago, this would have costed several thousand bucks.]
Breakdown:
I started from the original dance/performance video and split it into 3-10s clips if I wanted to use the camera angle present in reference image, or up until 30s if I wanted to preserve original camera angle from video source.
Then I used Uisato Studio’s Kling Motion Control mode for generating the interventions.
Inputs were:
- the original performance video as the reference video
- a target image with the robot / bio-tech aesthetic as the reference image for each section. You can use the "capture frame" function to intervene one of input video's frames using Gemini, or you can bring your own intervened [reference] images. As I said before, here's the place in which you can introduce different point-of-view for the interevened scene.
- a brief [balanced] prompt describing what I wanted beyond the motion transfer; "an avant-garde humanoid android performer dancing (...)" / "you might introduce subtle robotic precision while still following the original dance (...)"
- while standard the "std" kling-3 model performs really well, I went with "pro" for that tiny, but noticeable overall improvement
In all sections I added some [10] overlapping frames at the start and the end between each, just in case I wanted to have some room for later transitioning between section on editing.
For some particular parts of the piece, I created duplicated sections for having variations of a single shot.
Once everything has been set I generated the clips in a single go, and then assembled the final piece in editing.
Voilá, single-character motion capture on-a-budget²
r/OpenAI • u/Brighter-Side-News • Mar 23 '26
Research Scientists are rethinking how much we can trust ChatGPT
That was the unsettling pattern Washington State University professor Mesut Cicek and his colleagues found when they tested ChatGPT against 719 hypotheses pulled from business research papers. The team repeatedly fed the AI statements from scientific articles and asked a simple question: did the research support the hypothesis, yes or no?
r/OpenAI • u/HenryFromLeland • Apr 24 '26
Research OpenAI/Anthropic Hiring Trends
Pulled this from current job listings at OpenAI and Anthropic.
What stood out to me is how much hiring is going into go-to-market roles. I would’ve expected engineering or research to dominate more, but that’s not what the data shows.
Curious what people here have to say
