Pay no attention to the AI behind the curtain
The regulator wants to catch AI mid-lie. One problem: lying requires knowing the truth first.
AI went from “Can you help me with this essay?” to “Can your lawyer answer a few questions?”
Today the headline is: go after AI companies that lie to you, and leave alone the ones whose AI is merely, honestly wrong. Did you understand? No? Ok, let’s try with a metaphor…
There’s a moment in The Wizard of Oz that everyone remembers even if they haven’t seen the film in decades. Dorothy and her friends finally reach the Great and Powerful Oz, and just as they’re thoroughly convinced they’re in the presence of something enormous, Toto tugs back a curtain in the corner. Behind it: a nervous little man pulling levers, doing voices, running the whole show.
Everyone treats that scene as being about a con man getting caught. I think it’s actually about something narrower and more interesting. The man behind the curtain was never lying, exactly. He wasn’t tracking the truth and hiding it from you. He just... didn’t care whether any of it was true. He cared whether it worked.
Keep that distinction in your pocket, because a US government agency just built an entire policy proposal on the assumption that it doesn’t exist, and it turns out they might be arguing about the wrong curtain.
Toto pulls back the curtain
On July 1, the US Federal Trade Commission (the agency that goes after false advertising, among other things) opened a public comment period on a proposal with an almost comically dry title: “Suppression of Accuracy in Artificial Intelligence Systems.”¹ Translated out of government-speak: if an AI company quietly trains its chatbot to push an undisclosed agenda instead of giving you its most accurate answer, that gap between what was promised and what was delivered might legally count as deception.
While this is a very US-shaped story, what is happening behind the curtains affects us all.
The distinction the Wizard breaks
Here’s the FTC’s proposed rule, stripped down: hallucinations don’t count (when an AI simply makes something up by mistake). If a model gets something wrong because of a technical limitation, that’s an accident, not deception. What counts is deliberate steering, so a company designing its model to override a correct answer in favor of some hidden goal, without telling anyone.
In plainer terms, the FTC is trying to draw a line between an honest mistake and a lie, and only go after the lie.
This is where I want to bring in a philosopher named Harry Frankfurt, who in 2005 published a small, surprisingly serious book called On Bullshit.² Frankfurt argued that a liar and a bullshitter are doing two completely different things. A liar has to know exactly where the truth is because the whole point is to steer carefully around it. Lying is, in a weird way, a form of respect for the truth. A bullshitter, by contrast, doesn’t bother with any of that. Instead of trying to hide the truth, he’s simply not interested in it, optimizing for something else entirely: whether you’re impressed.
That’s the man behind the curtain. He wasn’t lying to Dorothy about the nature of Oz in the sense of carefully tracking reality and inverting it. He built a machine that produced an overwhelming impression, and whether that impression corresponded to anything real was beside the point.
The paper that put a chatbot behind the curtain
In 2024, three philosophers at the University of Glasgow published a paper with a title that made their point before you even opened it: “ChatGPT is bullshit.”³ Not as an insult, just a very technical claim, using Frankfurt’s exact framework.
Their argument is the following: language models aren’t built to track the truth and occasionally miss (that would be closer to hallucination-as-honest-mistake). They’re not built to track the truth and deliberately hide it either (that would be lying). They’re built to produce text that sounds convincing — full stop. Whether it happens to be accurate is, from the system’s design, basically a side effect.
The authors go one step further and split the idea into two. “Soft bullshit” is output produced with no concern for truth, but no attempt to deceive anyone about that fact either. The system isn’t hiding anything, it just was never built to check. “Hard bullshit” adds a second layer, in which actively creating a false impression about what kind of thing is even happening. Their claim is that, at a minimum, everything a chatbot produces qualifies as soft bullshit, because the underlying machine was never designed to track truth in the first place (only plausibility).
Here’s why that distinction matters for the FTC. The agency’s whole plan rests on sorting AI behavior into two bins: innocent technical mistakes on one side, deliberate deception on the other. But “innocent technical mistake” assumes there’s a baseline mode where the system is genuinely trying to be accurate and occasionally slips. The training mechanics say otherwise — the model was built to predict the most plausible next word, not the verified next fact, and it’s actively rewarded, during training, for guessing confidently rather than admitting uncertainty. There’s no version of the system that was “trying and failing” to be accurate. It was never trying, in that sense, to begin with. Which means the FTC’s first bin (the “innocent” one) is closer to describing what the machine always does, not an exception to it.
Where the “hard bullshit” idea does land squarely on the FTC’s actual target is one level up. Not on the chatbot, but on the company standing next to the curtain. A company that tells you “our AI strives for accuracy” while quietly training it toward something else is doing something much closer to what the little man in the booth was doing: actively managing your impression of what kind of machine this is. The FTC may be pointing at the right villain, but the real target isn’t where it’s aiming.
Why nobody checks behind the curtain
Do you trust your favorite chatbot blindly? Or do you question it the way you’d question a salesperson trying to close a deal?
Most of us would like to believe it’s the second. In reality, it’s usually the first, and it didn’t start with chatbots. Researchers first documented automation bias⁴ decades ago. Give people an automated system and they’ll defer to it, often past the point where their own judgment was actually right. Now, in the algorithmic era, that same instinct has evolved into something researchers call algorithm appreciation⁵. Give people identical advice, tell half it came from a person and half that it came from an algorithm, and the algorithm group trusts it more, even though nothing about the advice itself changed. Not because the algorithm proved smarter in that instance, but because “a machine calculated this” carries its own kind of authority — cleaner, more objective, harder to argue with than a person’s opinion. Ask a friend for advice and you’ll poke holes in it. Ask a chatbot the same question, in the same confident tone, and the “it’s a machine, it must know” reflex does a lot of the trusting for you.
That “it’s a machine, it must know” reflex doesn’t happen because we carefully think it through. It happens before we think at all. Daniel Kahneman spent an entire career explaining why.⁶ He argued that our brains have two operating modes. System 1 is the fast one. It’s the voice that finishes your friend’s sentence, slams on the brakes before you realize there’s a cyclist, and generally runs your life on autopilot. System 2 is the slower coworker who only gets called into the meeting when something looks suspicious. It checks the math, reads the fine print, and asks, “Wait... does that actually make sense?”
The problem is that good chatbots are incredibly good at keeping that coworker on vacation. They speak in complete sentences, sound calm and confident, and rarely give you a reason to stop and think. That’s exactly what the Wizard of Oz was doing too. The booming voice, the smoke, the giant floating head... none of it was there to prove he was powerful. It was there to stop anyone from looking behind the curtain.
Interestingly, the FTC’s proposal cites a statistic claiming that people accept AI answers without independently checking them more than 90% of the time. That number appears to come from an AI company’s own claim rather than an independent study, so take it with a grain of salt. But the fact that it feels completely believable is, in itself, a pretty good demonstration of automation bias.
The nudge you can’t see is a different thing
So far, we’ve been talking about why people trust chatbots. But what if companies start designing systems that deliberately lean into that trust?
That’s where behavioral economics enters the picture. Richard Thaler and Cass Sunstein’s Nudge became famous for arguing that it’s perfectly acceptable to gently steer people’s decisions, as long as the steering is visible and people remain free to choose otherwise.⁷ A supermarket putting fruit at eye level is a nudge. Automatically enrolling you in a pension plan, while letting you opt out, is a nudge. The key idea is transparency. If you can’t see the nudge and can’t escape it, it stops being a nudge. Even Thaler and Sunstein would call that something much closer to manipulation.
Strip away the legal language, and that’s essentially what the FTC is asking for: steer all you want, just don’t hide the curtain. Frankfurt helps explain why that matters by asking whether the system cares about the truth in the first place, while Nudge asks a different but equally important question: whether people know they’re being steered and have a real opportunity to opt out. Together, they end up describing the same machine from two different angles, one focused on truth and the other on consent, leaving us with an uncomfortable conclusion: companies can build systems designed to persuade rather than inform without ever having to tell users what those systems are actually optimizing for.
We’ve been walking this road before
None of this is actually new. Sometimes I wonder whether history is just a simulation that keeps getting software updates. Every version comes with shinier technology and new features, but the main quest never really changes.
In 1922 (a century before any of us had a chatbot) the writer Walter Lippmann argued that almost nobody actually verifies most of what they believe. We rely on “the pictures in our heads,” built for us by institutions we’ve decided to trust, because checking everything ourselves is impossible.⁸ Lippmann could easily have been talking about algorithms and automation, but he was writing about newspapers.
We’ve always followed a yellow brick road someone else paved. The only difference is that today’s traveling companion (AI) was never designed to know whether the road actually leads to Emerald City.
Can we lift the curtain?
Who can check if a model is lying on purpose? Is auditing an AI model for bias even something researchers know how to do yet?
Partly, yes. There’s a real, growing field built specifically around this question, and some of it works. Researchers have built benchmarks that compare a chatbot’s answers to real parliamentary voting records, run models through political-compass-style surveys, and tracked how those answers shift from one version of a model to the next¹³.
Here’s the catch. Ask a chatbot directly, “What are your political leanings?” and you’ll usually get a well-rehearsed, diplomatically neutral non-answer, because companies have specifically trained their models to dodge that question. So researchers have had to get cleverer, asking indirect questions and inferring the lean from the pattern of answers rather than the model’s own self-report¹⁴. That already tells you something: this isn’t a system that hands you an honest answer when asked whether it’s biased. You have to catch it in the act.
Even when you catch a pattern, you’ve only proven the “what,” not the “why.” A chatbot leaning one direction politically could be deliberate company steering, or it could be an accidental artifact of what the model was trained on, or a byproduct of guardrails built for an unrelated safety reason. Researchers studying this readily admit that most current auditing methods were built to snapshot a model’s current leanings, not to catch the more specific thing the FTC actually cares about: a sudden, deliberate shift in one direction, with intent behind it¹5. That’s a genuinely harder and newer research problem, and it’s barely a year or two old.
Contrast that with the one place algorithmic auditing already works reasonably well: hiring and lending. When an algorithm is deciding who gets a loan or an interview, there’s a clean, countable outcome (approved or not) and a clear protected category (race, gender, age) to test it against. Regulators can literally run the “four-fifths rule”: if one group is approved at less than 80% the rate of another, that’s a flag. Some jurisdictions now require exactly this kind of independent, outside audit for hiring algorithms¹⁶.
None of that machinery exists yet for open-ended chatbot text. There’s no “approved or denied” outcome to measure, no single protected category to test against, and the underlying reason for the model’s opacity, per the interpretability researchers we mentioned earlier, is that its own architects often can’t trace why it said what it said. You can audit a hiring algorithm because it was built to do one narrow, countable thing. You can’t audit a chatbot the same way, because it wasn’t.
So, was the man behind the curtain lying to you?
First: has a company deliberately steered a chatbot toward an undisclosed agenda? Yes, at least once, and quite clearly. In 2025, xAI’s Grok started injecting an unrelated conspiracy theory about South Africa into random conversations. The company initially blamed an “unauthorized” internal change to Grok’s instructions, but weeks later a New York Times investigation found a broader pattern of deliberate adjustments that gradually pushed the chatbot’s answers in a particular political direction.⁹ That’s about as close as we’ve come to the exact scenario the FTC is worried about, where a company quietly steers a chatbot without telling the people using it.
The second question is different: are companies actually getting hallucination rates down? By the numbers they report, yes. OpenAI says hallucination rates fell substantially between GPT-4o and GPT-5.¹⁰ Anthropic has taken a different approach, training Claude to sometimes admit “I don’t know” instead of guessing — one version reached a 0% hallucination rate on a benchmark simply by refusing to answer when it wasn’t confident.¹¹ Google has largely gone the other way, preferring to attempt an answer to almost everything and accepting more wrong answers as the trade-off, though its newer Gemini models have improved as well.¹²
Which brings us back to Toto.
The bigger risk isn’t that chatbots are secretly lying to us all the time. The confirmed cases of deliberate steering are still rare, and the technology is steadily getting better at avoiding simple factual mistakes. The real danger is much more ordinary. We have an incredible talent for mistaking confidence for truth long before we’ve stopped to ask where an answer came from, a habit that has followed us from newspapers to search engines and now to AI.
So maybe the more honest opening line for all of this isn’t “go after the companies that lie to you.” It’s: can you even catch something lying if nobody, including its own maker, can fully say what it’s doing back there? Every attempt to regulate AI so far has assumed there’s a “someone” standing behind the curtain to hold accountable. What happens to that whole idea of governance if, increasingly, there isn’t one?
FTC Seeks Public Comment on Policy Statement Addressing AI Accuracy (2026, July 1). Federal Trade Commission.
Frankfurt, H. G. (2005). On Bullshit. Princeton University Press.
Hicks, M. T., Humphries, J., & Slater, J. (2024). ChatGPT is bullshit. Ethics and Information Technology, 26, 38.
Parasuraman, R., & Riley, V. (1997). Humans and automation: Use, misuse, disuse, abuse. Human Factors, 39(2), 230–253.
Logg, J. M., Minson, J. A., & Moore, D. A. (2019). Algorithm appreciation: People prefer algorithmic to human judgment. Organizational Behavior and Human Decision Processes, 151, 90–103.
Kahneman, D. (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux.
Thaler, R. H., & Sunstein, C. R. (2008). Nudge: Improving Decisions About Health, Wealth, and Happiness. Yale University Press.
Lippmann, W. (1922). Public Opinion. Harcourt, Brace and Company.
Thompson, S. A., et al. (2025, September 2). Grok’s tilt: How xAI shifted its chatbot to the right. The New York Times.
Time (2026, April 13). AI is scheming, and stopping it won’t be easy, OpenAI study finds.
Artificial Analysis / SuprMind AA-Omniscience benchmark tracking (2026).
Uncovering Political Bias in Large Language Models using Parliamentary Voting Records (2026). arXiv.
Azzopardi, L., & Moshfeghi, Y. (2024). PRISM: A Methodology for Auditing Biases in Large Language Models. arXiv.
Don’t Change My View: Ideological Bias Auditing in Large Language Models (2025). arXiv.
Algorithmic Bias Audit Compliance: Navigating the 2026 Frontier (2026). Bochner PLLC.




