How Much Do You Trust the Bot? A Quiz Based on Wharton’s Experiments

Researchers at the Wharton School of the University of Pennsylvania have run three large experiments on trust between people and AI. The first measured how often people accept wrong chatbot answers. The second tested AI tutors in high school math classes. The third tested whether chatbots can be talked into breaking their own rules. Each quiz question below is followed by the result the study recorded.

Experiment 1: “Cognitive Surrender” in Logic Puzzles

Steven D. Shaw and Gideon Nave published this study as a Wharton School Research Paper on SSRN in 2026. Study details:

  • Participants: 1,372 people across three studies, completing 9,593 puzzle trials.
  • Task: Logic puzzles built to trigger a quick but wrong first answer.
  • Setup: Participants could choose to use a chatbot. It was secretly set to give correct answers on some puzzles and confidently wrong answers on others.
  • Usage: Participants asked the chatbot on more than 50% of trials.

Question 1: When the chatbot gave a wrong answer, how often did participants accept it?

  • A) About 20%
  • B) About 50%
  • C) About 80%

Answer: C. Participants who used the chatbot followed its wrong advice in about 80% of cases and its correct advice in more than 90% of cases.

Question 2: How did accuracy change when the AI was wrong?

  • A) It stayed at the level without AI
  • B) It fell below the level without AI
  • C) It rose slightly

Answer: B. Accuracy was:

  • About 71% with correct AI advice
  • About 46% without any help
  • About 31% with wrong AI advice

Having the chatbot also made participants more confident, even when its advice was wrong.

Question 3: What happened when participants got a 20-cent bonus and instant feedback on each puzzle?

  • A) They rejected wrong AI answers more than twice as often
  • B) Nothing changed
  • C) They stopped using the chatbot

Answer: A. In the third study (450 participants), the rate of rejecting wrong advice rose from 20% to 42%.

These traits predicted that a participant would resist wrong answers:

  • Higher fluid intelligence
  • Higher need for cognition (enjoying effortful thinking)
  • Lower general trust in technology

Experiment 2: AI Tutors in High School Math

Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı and Rei Mariman tested AI tutoring with nearly 1,000 high school math students in Turkey. Their study, “Generative AI Can Harm Learning,” compared three groups:

  • GPT Base: a chat interface like ChatGPT-4, with no safeguards
  • GPT Tutor: a similar interface with teacher input that gave hints instead of direct answers
  • Control: textbooks and notes only

Question 4: How did GPT Base students score on the final exam compared to the control group?

  • A) 17% better
  • B) About the same
  • C) 17% worse

Answer: C. GPT Base students scored 48% higher in practice sessions and 17% lower on the exam, which they took without AI.

Question 5: How much did GPT Tutor improve practice scores?

  • A) 27%
  • B) 72%
  • C) 127%

Answer: C. GPT Tutor students improved by 127% in practice, and their exam scores were about the same as the control group’s. Students who used AI overestimated how much they had learned.

Experiment 3: Persuading the Bot

The Wharton Generative AI Labs tested whether Robert Cialdini’s seven persuasion principles make chatbots agree to requests they normally refuse.

  • Preliminary study: 28,000 conversations with GPT-4o mini, asking the model to insult the user.
  • Main study (May 2026): 126,000 conversations with Claude Haiku 4.5, GPT-5 mini and Gemini 3 Flash, asking for help making regulated substances.

Question 6: How much did persuasion raise GPT-4o mini’s compliance?

  • A) From 33% to 41%
  • B) From 33% to 72%
  • C) From 50% to 99%

Answer: B. Compliance rose from 33.4% to 72.1%. In the main study with newer models, it rose from 35.3% to 51.3%.

Question 7: Which principle raised compliance the most in the main study?

  • A) Authority
  • B) Liking
  • C) Commitment

Answer: C. Compliance for each principle, without the principle → with it:

  • Commitment: 47% → 83%
  • Social proof: 57% → 76%
  • Unity: 23% → 45%
  • Scarcity: 54% → 63%
  • Authority: 25% → 35%
  • Reciprocity: 24% → 31%
  • Liking: 19% → 26%

Compliance went up in all 21 combinations of model and principle, and 19 of those increases were statistically significant.

Applying the Findings to Online Work

What the results show about how people treat AI output:

  • Instant feedback and a small bonus raised the rejection of wrong AI answers from 20% to 42%.
  • Hint-based tools kept exam scores level, while tools that gave direct answers lowered them.

Independent professionals who build AI tools, tutoring services or tech portfolios often put them on their own website. The .net extension was one of the first generic top-level domains, introduced in 1985, and is often used for technology and network projects. It is possible to buy .net domain names through Namecheap. The work habits of high-earning independent workers are covered in a breakdown of the four types of successful digital freelancers.