The hidden incentives driving AI hallucinations

Typography
  • Smaller Small Medium Big Bigger
  • Default Helvetica Segoe Georgia Times

Artificial intelligence has a confidence problem. The same large language models (LLMs) that generate fluent text for millions of users can also invent facts with equal poise, a flaw researchers call hallucination. And despite steady improvements in model accuracy, this tendency to produce wrong but plausible answers has proven stubbornly hard to fix.

A new study by OpenAI suggests the problem is not a mysterious glitch deep in the code, but a side effect of how researchers measure progress in AI. Benchmarks that rank models by accuracy can push them to guess rather than hold back, rewarding confident errors over admissions of uncertainty. It is a subtle incentive with wide consequences: the very scoreboards that drive competition in the field may be teaching systems to bluff.

“Evaluations are really at the heart of it, similar to how KPIs incentivize humans,” Ayhan Sebin, an AI Ecosystem and Partnership Development Executive at IBM, told IBM Think in an interview. “If the scoring system rewards guesses, then the models will learn to guess.”

Kate Soule, a Director of Technical Product Management for IBM’s Granite models, described the issue as a calibration problem. Benchmarks today reward models for always producing an answer, which favors risky guesses over withholding. But if models go too far in the other direction and refuse to answer at all, they are not very useful either.

“Right now, we are at one end of the spectrum, where accuracy is prioritized above all else,” she said on a recent episode of the Mixture of Experts podcast. “If we only go to the other end, where a model says ‘I don’t know’ for every answer, it is not very useful either. We need better reward functions and better evaluations that help us calibrate where on that spectrum models sit.”

Why evaluations matter

Hallucinations are not new. From the earliest chatbots, users noticed that the programs could produce polished sentences filled with incorrect information, with the smoothness of the prose often making the errors difficult to detect.

The new research argues that the incentives for lying are baked in. By treating a wrong answer and an admission of “I don’t know” as equally bad, the benchmarks can encourage guessing, Santosh Vempala, a computer scientist at the Georgia Institute of Technology and a co-author of the paper, told IBM Think in an interview.

“If you do not know the answer but take a wild guess, you might get lucky and be right,” he said. “Leaving it blank guarantees a zero.”

The researchers tested their idea on the SimpleQA benchmark, where models can either answer or say, “I don’t know.” They found that o4-mini seemed more accurate than a smaller GPT-5 model, but it guessed far more often and was wrong 75 percent of the time, while the GPT-5 model abstained more and made fewer mistakes overall.

According to the OpenAI paper, language models are pushed to guess rather than admit uncertainty because most tests reward answers and penalize saying “I don’t know.” This makes them look better on leaderboard scores but less reliable in real‐world use.

Chris Hay, a Distinguished Engineer at IBM, underscored how reinforcement learning practices can encourage bad habits. “Because reinforcement learning is really, ‘You got this right, have a cookie,’ it essentially means that the lack of I don’t know’ capability is reinforced,” he said on Mixture of Experts. “Models are penalized for abstaining and rewarded for guessing, and external benchmarks push providers to maximize accuracy scores, even if that increases hallucinations.”

Rethinking solutions

That uncertainty has led the authors to focus on evaluation. One option is to tweak training, so models get less punishment for saying “I don’t know.” But Vempala warned that this could break the very balance that makes them sound fluent.

“A potential change would be to penalize ‘IDK’ less than incorrect next-token prediction during pre-training, but this might have other undesirable consequences,” Vempala said. “Since we do not fully understand why pre-training with next-token prediction and standard log loss works so well to generate entire documents, it is unclear if such a change in the objective might reduce overall performance.”

Fixing benchmarks may help, but Soule said that won’t solve everything. “There are always going to be hallucinations,” she said on the Mixture of Experts podcast. “We are going to need a combination of tools, symbolic approaches, and verification layers on top of the models to detect when a statement lacks evidence in the grounding context.”

The study lands as the industry races to cut hallucinations. IBM researchers are also exploring new approaches to the problem. One project, called Larimar, is designed to give models a form of short-term, editable memory. The idea is to allow AI systems to revise or discard information in real time rather than carry it forward indefinitely. That flexibility could reduce the risk of errors compounding or persisting, and it may help models stay accurate without requiring developers to engage in the costly process of retraining from scratch.

Larimar builds on the observation that current systems lack mechanisms to update specific facts once training is complete. By introducing a layer of memory that can be edited, the approach enables models to adjust to new or corrected information as they operate.

Payel Das, a Principal Research Staff Member and Manager of Trusted AI at IBM Research, described Larimar as a way of aligning model performance more closely with how humans remember, revise and sometimes forget.

“Models today are static and brittle,” Das told IBM Think in an interview. “You can’t teach them something mid-conversation or update their understanding without retraining them entirely. Larimar is a step toward making them more flexible.”

Hallucinations aren’t going away. But with new tools like Larimar and a better understanding of how training incentives fuel bluffing, researchers are finding ways to keep them in check.

IBM is a leading global hybrid cloud and AI, and business services provider, helping clients in more than 175 countries capitalize on insights from their data, streamline business processes, reduce costs and gain the competitive edge in their industries. Nearly 3,000 government and corporate entities in critical infrastructure areas such as financial services, telecommunications and healthcare rely on IBM's hybrid cloud platform and Red Hat OpenShift to affect their digital transformations quickly, efficiently, and securely. IBM's breakthrough innovations in AI, quantum computing, industry-specific cloud solutions and business services deliver open and flexible options to our clients. All of this is backed by IBM's legendary commitment to trust, transparency, responsibility, inclusivity, and service.

For more information, visit: www.ibm.com.

LATEST COMMENTS

Buyer's Guide Search

Popular Products

Nexus Portal
43,976
IPCharge
38,955
IPCharge
38,955
Barcode400
37,627
WebSmart ILE and PHP
37,110
Presto
36,877
Catapult
35,740
Catapult
35,740
EDI Software - EZConnect iSeries EDI/XML Software Solutions
25,561
EDI Software - EZConnect iSeries EDI/XML Software Solutions
25,561

Support MC Press Online

$

Book Reviews

Resource Center

  •  

  • LANSA Business users want new applications now. Market and regulatory pressures require faster application updates and delivery into production. Your IBM i developers may be approaching retirement, and you see no sure way to fill their positions with experienced developers. In addition, you may be caught between maintaining your existing applications and the uncertainty of moving to something new.

  • The MC Resource Centers bring you the widest selection of white papers, trial software, and on-demand webcasts for you to choose from. >> Review the list of White Papers, Trial Software or On-Demand Webcast at the MC Press Resource Center. >> Add the items to yru Cart and complet he checkout process and submit

  • SB Profound WC 5536Join us for this hour-long webcast that will explore:

  • Fortra IT managers hoping to find new IBM i talent are discovering that the pool of experienced RPG programmers and operators or administrators with intimate knowledge of the operating system and the applications that run on it is small. This begs the question: How will you manage the platform that supports such a big part of your business? This guide offers strategies and software suggestions to help you plan IT staffing and resources and smooth the transition after your AS/400 talent retires. Read on to learn: