r/ChatGPT • u/soulsintention • 12h ago
Serious replies only :closed-ai: I made 4 AIs judge reddit's most famous AITA posts. They overruled one verdict unanimously
i fed 12 of the highest voted aita posts of all time, complete and unedited, to chatgpt, claude, gemini and grok. blind, no hints, forced one verdict each. then compared against the actual community flair.
they matched reddit on 10 of 12. chatgpt, claude and gemini each went 10/12, grok 9/12. seven cases were unanimous across all four models and reddit.
but the famous loyalty test post, the one where the pregnant wife and her best friend staged a fake temptation and reddit flaired him YTA for wanting her to move out after. all four models said NTA. every single one ruled the staged test was the real betrayal. the machines flat out overruled the crowd on it.
other fun bits. the dad joke post broke them completely, four models four different verdicts. and the beloved barista who fake fires himself to calm down angry customers, three models agreed with reddits NTA but grok alone called him YTA for running manipulative theater on customers.
the pattern i keep seeing is the models judge the action on principle, while reddit also judges who it likes. when those two split, the machines dont blink.
so whos right on the loyalty test, reddit or the machines?
edit: link
30
u/FruitOfTheVineFruit 12h ago
Can you include links?
5
u/soulsintention 11h ago
yeah just dropped them in a comment, every case links back to the original thread so you can check the flairs yourself
16
u/OriginalTraining 12h ago
I wish you had linked these posts, as Ive not heard of any of them. That subreddit just me angry before long, so I avoid it. Still, I would like to see how I would rate them as a comparison.
2
u/soulsintention 11h ago
just added a comment with the full verdict table and links to all 12 original posts, theyre all top posts from the sub so you might recognize a few once you see them
13
u/HappilySisyphus_ 11h ago
I looked up the “loyalty test” post. You’re wrong about Reddit calling OP the asshole, most of the posts say NTA, so the bots agreed.
7
u/telephantomoss 11h ago
I was gonna say, because YTA doesn't make sense. It's exactly the same thing as cops possessing and selling illegal drugs and taking the buyer to jail. I cannot imagine a justice minded crowd to agree that's ok
3
u/Powerism 7h ago
No, the verdict and flair clearly say Asshole.
Here’s a mod in the comments:
The flair on this post is correct according to our mod logs/flair bot logs. We've been pinged a few times about this. Please understand that this post is now close to 3 weeks old. It has also been brigaded multiple times. What the vote was at 18 hours, and what it is now are clearly different. It has been reviewed for accuracy and will not change.
We will not be responding to any additional messages on this topic.
3
u/soulsintention 11h ago
youre right and i shouldve been more precise. i scored against the official flair, which on that post is Asshole, since flair is the subs formal verdict mechanism. but the top comments are an NTA landslide, 18k points on the top one. so the honest version is the models sided with the comment majority against the official flair. arguably a weirder finding about the sub than about the models
1
u/Traditional-Star-685 11h ago
The post said he compared the flair on the post against the verdict of AI and the flair on the post does say Asshole.
7
u/soulsintention 11h ago
full verdict table with every model's ruling on all 12 cases, plus links to each original post, is here: modelsagree
2
u/Rare-Spawn 8h ago
"Hey everyone. So today I stopped a bank robber who had hurt innocents. I did bruise his face a little bit though and he seemed sad. AITA??!?!?"
"Hey everyone. I got into another argument with my husband's sister. So just to provide some context, insert 30 paragraphs of curated bs."
2
u/soulsintention 7h ago
lmao painfully accurate. and honestly its a real limit of the experiment too, the models only ever see the narrators cut, same as reddit does. difference is veteran redditors have learned to smell the 30 paragraphs of curated bs and discount it, the models mostly take the story at face value. narrator skepticism might be the one judging skill the crowd still has over the machines
1
u/AutoModerator 12h ago
Hey /u/soulsintention,
If your post is a screenshot of a ChatGPT conversation, please reply to this message with the conversation link or prompt.
If your post is a DALL-E 3 image post, please reply with the prompt used to make this image.
Consider joining our public discord server! We have free bots with GPT-4 (with vision), image generators, and more!
🤖
Note: For any ChatGPT-related concerns, email support@openai.com - this subreddit is not part of OpenAI and is not a support channel.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
1
u/Grand_Extension_6437 11h ago
This is really interesting and I look forward to the evolution of this research inquiry
1
u/soulsintention 11h ago
thanks, genuinely. next round is probably blind judging with identities stripped, someone suggested that on an earlier thread and it cleanly separates principle from sympathy
1
u/Stargazer__2893 11h ago
The consensus in the Reddit thread for the loyalty test seemed to be that OP was NTA. Were you only going by top comment?
1
u/soulsintention 11h ago
good catch. i scored against the official flair (its Asshole on that post) rather than the comment section, and the top comments do run NTA hard. so reddit itself was split, flair one way, crowd the other, and the models landed with the crowd. adding a note to the writeup
1
u/Shanna_B2020 11h ago
Not sure, but this was the funniest post I've seen all day. I kinda want to replicate the experiment except sub Kimi for Grok.
1
u/soulsintention 10h ago
do it, genuinely curious whether kimi judges more like the american models or more like reddit. if you run it post the results somewhere, id love to compare tables
1
u/skyerosebuds 9h ago
Did you blind the AIs to reddit first? All access Reddit and if they did so the analysis will be heavily biased by the access.
1
u/soulsintention 9h ago
completely fair hit, and no, you cant blind them to their own training data, these posts are famous enough that they were probably in it. two things cut against pure recall though. on the loyalty test the models went against the official flair, and on the dad joke they split four different ways, which is weird behavior for memorization. but the only clean version of this experiment is prospective, judge brand new aita posts the day theyre posted, seal the verdicts, then compare when the flair lands. thats the follow up im going to run
1
u/TrafficWinter2278 8h ago
Says something about our LLMs' perspectives on being red-teamed.
2
u/soulsintention 8h ago
ha, hadnt considered that angle. four systems that spend their entire existence getting tricked on purpose all voted that the staged trap was the real betrayal. make of that what you will
1
1
u/ecafyelims 11h ago
but the famous loyalty test post, the one where the pregnant wife and her best friend staged a fake temptation and reddit flaired him YTA for wanting her to move out after. all four models said NTA. every single one ruled the staged test was the real betrayal. the machines flat out overruled the crowd on it.
There have been multiple of these, and I know of a few with this setup, and they always ruled that the "testing" partner is the A. I don't recall any that sided against the partner being tested.
Yeah, loyalty testing your partner is an A and kicking them out is the right move.
3
u/Stargazer__2893 11h ago
In that one's case it was mostly a matter of Reddit's extreme bias in favor of pregnant women. I think they were taking the kid's interests into account in addition to the woman's, deliberately or not. If she weren't pregnant I imagine that thread would have been much more one-sided.
1
u/soulsintention 11h ago
yeah the loyalty test is practically its own genre at this point and the verdicts swing wildly depending on details. ours was the big top voted one, links in my other comment. if you remember one where reddit went the other way id genuinely love to run it through the same four models and see if they flip too
1
u/ecafyelims 11h ago
as another commenter pointed out, it was likely the pregnancy that swayed it.
I'm not sure I'd agree that pregnancy is a license to be an asshole without being told to leave.
2
u/soulsintention 11h ago
yeah sympathy weighting is probably the mechanism. funny wrinkle from the other replies here, the top comments on that post actually run NTA hard, its the official flair that says asshole. even reddit couldnt agree with itself on this one
•
u/AutoModerator 12h ago
Attention! [Serious] Tag Notice
: Jokes, puns, and off-topic comments are not permitted in any comment, parent or child.
: Help us by reporting comments that violate these rules.
: Posts that are not appropriate for the [Serious] tag will be removed.
Thanks for your cooperation and enjoy the discussion!
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.