For Immediate Release
A lessons-learned digest tracing how the distance between persuasive words and verified truth shows up in both AI systems and education policy and what the convergence reveals about building more trustworthy systems.
Artificial intelligence, despite its impressive ability to generate human-like text, fundamentally lacks mathematical reasoning skills, creating a significant and potentially dangerous honesty gap. This weakness means AI can present incorrect information with convincing fluency, making it difficult for users to discern truth from falsehood. The illusion of understanding, combined with a lack of genuine comprehension, poses a growing risk as AI becomes more integrated into daily life.
This is the central tension explored in The Honesty Gap: Words Vs. Math from GenXis Research, a new framework for understanding why artificial intelligence systems can be simultaneously impressive and unreliable. The anxiety around AI, as the research explains, is not merely that machines can be wrong. It is that machines can be wrong in fluent, reasonable, socially persuasive language.
The term "honesty gap" didn't originate in technology discussions. It emerged from education policy, where it described something more tangible: the measurable distance between how students perform on state-administered tests versus the National Assessment of Educational Progress (NAEP), often called the Nation's Report Card. When states lowered their proficiency thresholds, achievement data painted an optimistic picture one that affected funding, reputation, and political narrative.
The U.S. Chamber of Commerce Foundation's April 2026 brief on the topic frames it clearly: the honesty gap measures the difference between how students perform on the national gold-standard assessment and how they perform on their own state's tests. The brief, produced in partnership with the Collaborative for Student Success, documents a pattern that spans years of data.
In 2024, Iowa's state-reported eighth-grade math proficiency rate stood at 72%. NAEP reported only 27% a 45-percentage point difference. Virginia's 2024 state-reported fourth-grade reading proficiency rate was 73%. NAEP reported 31% a 42-point gap. These aren't edge cases. They represent systematic misalignment in how "proficiency" is defined and measured.
"If we believe that NAEP is indeed the Nation's Report Record on student proficiency, then we would hope there is little difference between the outcomes on the two tests," said Jim Cowen, Executive Director of The Collaborative for Student Success, in a 2024 analysis of the honesty gap. "But that's not the case. In many states, the gaps suggest that parents simply aren't getting the full picture of how prepared their kids are for college or the workforce."
What does education policy have to do with artificial intelligence? More than might first appear. Both domains grapple with the same fundamental problem: the distance between persuasive language and verified truth.
The GenXis Research framework calls this the Honesty Gap, and it extends the education concept into a broader analysis of how language metabolizes error. "Words can escape meaning," the research explains. "They can rationalize, soften, blur, excuse, reframe, and drift." In human psychology, this shows up as motivated reasoning, cognitive dissonance reduction, and what researchers call ethical fading the gradual disappearance of ethical considerations from verbal reasoning.
In AI systems, the same dynamic appears as hallucination: the generation of unsupported synthesis and citation-shaped language without source custody. A legal citation can be fabricated in perfect legal prose. A medical explanation can sound clinically plausible while omitting a contraindication. A financial summary can appear authoritative while relying on stale facts.
"A legal citation can be fabricated in perfect legal prose. A medical explanation can sound clinically plausible while omitting a contraindication. A financial summary can appear authoritative while relying on stale facts."
The danger comes from the mismatch between linguistic confidence and verified grounding. Natural language is flexible by design. It allows approximation, metaphor, implication, emphasis, ambiguity, and context dependence. Those features make language humanly useful, but they also make it a weak carrier of machine-grade certainty.
Cory Koedel, a tenured professor of economics and public policy at the University of Missouri-Columbia, has studied the honesty gap in education for more than two decades. Writing in the Show-Me Institute's analysis published April 18, 2025, Koedel describes how the education system "often fails to communicate honestly with students, parents, and community members about how much students are actually learning."
The numbers are stark. Approximately 90% of parents believe their children are performing at or above grade level in reading and math. Only about one-third of fourth- and eighth-grade students in the United States score at a proficient level on NAEP. "Grades are up, but test scores are down," Koedel notes. "This is problematic because grades tend to carry more weight with students and parents than test scores."
The parallel to AI is direct. Users of language models often assign disproportionate trust to outputs simply because they arrive in confident, well-formed sentences. The linguistic polish creates an illusion of reliability.
The education honesty gap provides concrete, quantifiable examples of how language standards can diverge from verified outcomes. The U.S. Chamber Foundation's state-by-state analysis for 2023-2024 reveals patterns that illuminate the broader problem.

| State | Grade | Subject | State Test Proficiency | NAEP Proficiency | Gap (Percentage Points) |
|---|---|---|---|---|---|
| Alabama | 4 | Reading/ELA | 58% | 28% | -30% |
| Iowa | 8 | Math | 72% | 27% | -45% |
| Virginia | 4 | Reading | 73% | 31% | -42% |
| Michigan | 8 | Reading | 65% | 24% | -41% |
| New York | 4 | Math | 50%+ | <40% | -10%+ |
Source: U.S. Chamber of Commerce Foundation Honesty Gap Brief (2023-2024 data)
In Virginia specifically, the Thomas Jefferson Institute documented how the state's Standards of Learning (SOL) assessment defines "proficient" in reading at a level that aligns to "below basic" on the national assessment. That means students who fail to display even partial mastery of knowledge and skills fundamental for grade-level work are deemed proficient by state standards. Virginia is one of only two states where this misalignment occurs.
"You will hear that NAEP 'proficient' is too high a bar and not a good proxy for the ability to read with comprehension," said Robert Pondiscio, senior fellow at American Enterprise Institute, in the Thomas Jefferson Institute's analysis. "A fair point as far as it goes, but I defy you to find me a single parent comfortable with her child reading at 'below basic' level."
The Fordham Institute's February 2025 commentary on the honesty gap traces how the problem has evolved. Dale Chu's analysis in Mind the honesty gap notes that the latest NAEP results were "disastrous," and the longstanding honesty gap issue is "no longer just a policy concern it's a glaring failure of accountability."
The commentary highlights compounding problems: "In New York, over half of fourth graders were deemed proficient in math on the state test in 2024 compared to less than 40 percent on NAEP. In Michigan, the gap is even starker: 65 percent of eighth graders were proficient in reading according to the state exam, while just 24 percent cleared the same bar on the nation's report card."
Grade inflation, the Fordham analysis notes, "only got worse during the pandemic and has since widened both performance and attendance gaps." The disconnect between perception and reality stymies progress and makes it harder to ensure students are truly prepared.
The Fordham commentary references a warning issued more than fifteen years earlier: in their introduction to The Proficiency Illusion, Checker Finn and Mike Petrilli described the same dynamics, noting that the testing infrastructure "on which so many school reform efforts rest, and in which so much confidence has been vested, is unreliable at best."
The education honesty gap offers a useful model for understanding what happens inside large language models. In both cases, the gap emerges from the same root cause: language standards that drift from verified outcomes.
In education, political pressures incentivize states to lower proficiency thresholds. Inflated rates mask learning gaps, misallocate resources, and leave students unprepared for the workforce. The U.S. Chamber Foundation brief explicitly notes that "declining student achievement limits the talent pipeline and threatens economic vitality."
In AI, no explicit political pressure exists but similar dynamics emerge from training objectives. Models optimized for fluency, coherence, and engaging prose will generate outputs that sound right, regardless of whether they are. The optimization target is linguistic quality, not verified accuracy.
GenXis Research frames it precisely: the central question becomes "when does a sentence become a verified claim?" Their Definition 1 provides formal grounding: "A claim is not merely a sentence. It is a tuple where is the statement, the domain, the truth condition, and the evidence requirement. Without those elements, language remains expressive but under-bounded."
The education honesty gap has been documented and debated for more than fifteen years. That history offers a lessons-learned digest for anyone working on AI reliability. The key insights are concrete:
Lowering standards doesn't improve outcomes. Virginia's experience demonstrates this clearly. Despite SOL proficiency rates that suggest strong performance, 42% of Virginia's fourth graders and 34% of eighth graders were reading below basic level on the 2024 national assessment. More than one-in-three Virginia students could not show even partial mastery of reading skills necessary for grade-level proficiency. Pretending otherwise doesn't change reality it just delays necessary intervention.
Parents deserve accurate data. The Collaborative for Student Success analysis frames this as a matter of trust: "The truth matters." States that embrace transparency rather than masking it demonstrate that accurate reporting, while uncomfortable, enables better decision-making. Massachusetts and Rhode Island closed their gaps to within five percentage points or less across both grades and subjects proving that alignment is achievable.
Calibrated measurement beats optimistic estimation. The Common Core and its associated exams significantly narrowed education honesty gaps when implemented, the Fordham analysis notes. Rigorous, consistent standards create reliable data. The same principle applies to AI: systems that ground outputs in verified constraints will generate more trustworthy results than those optimized for fluency alone.
GenXis Research's framework offers a direct answer to the honesty gap problem. The antidote, they argue, is "not less language, but stronger grounding." Specifically:
These approaches don't eliminate language they augment it. The goal is not to make AI systems less expressive, but to ensure that expressions carry verifiable weight.
The education honesty gap data isn't uniformly grim. The Collaborative for Student Success analysis identifies meaningful progress: "Massachusetts and Rhode Island closed their gaps to within 5 percentage points or less across both grades and subjects." Fourteen states are holding students to an equal or higher standard than NAEP in at least one grade or subject.
More significantly, states have improved over time. In 2014, 23 states had "the biggest honesty gaps" in fourth-grade reading defined as 30 percentage points or larger. By 2024, only Alabama, Iowa, Nebraska, and Virginia had gaps that large. In eighth-grade math, 14 states had gaps that large in 2014; by 2024, only Iowa, Mississippi, and Virginia remained in that category.
Virginia has also committed publicly and explicitly to addressing the problem. The state redesigned its school accountability and accreditation system and committed significant funding to high-dosage tutoring and literacy initiatives. The truth matters, and Virginia is now facing it directly.
For readers researching AI systems, the lessons from education policy are directly applicable. The honesty gap isn't a new problem it has existed in human institutions long before artificial intelligence emerged. The education sector's experience demonstrates both the dangers of unchecked optimism and the feasibility of rigorous alternatives.
When evaluating AI systems for high-stakes applications legal drafting, medical triage, financial analysis, scientific writing look for evidence of mathematical grounding. Ask whether outputs include source citations with verifiable custody chains. Note whether the system explicitly declines uncertain claims rather than hedging with confident-sounding approximations. These aren't minor technical details; they are the difference between a system that can be trusted and one that merely sounds trustworthy.
The fluency trap is real. Language that sounds right can be wrong in ways that are difficult to detect without systematic verification. This is true for AI outputs, and it is equally true for the reports, dashboards, and summaries that organizations use to make decisions. The education honesty gap offers a cautionary tale: optimistic data doesn't create optimistic outcomes. Only verified improvement does that.
For the full framework on AI and the honesty gap, explore GenXis Research's The Honesty Gap: Words Vs. Math, which provides the formal definitions and technical analysis underlying this issue.
For state-by-state education data, the U.S. Chamber of Commerce Foundation's April 2026 Honesty Gap Brief offers the most comprehensive current dataset comparing NAEP results to state proficiency rates.
For policy analysis and commentary, Fordham Institute's Mind the honesty gap traces the education policy implications and historical context of the accountability failure.
For Virginia-specific data and institutional response, the Thomas Jefferson Institute's analysis documents the concrete consequences of misaligned standards.
For the latest national analysis and bright spots, see the Collaborative for Student Success's 2024 Honesty Gap Analysis, which includes state-by-state comparisons and expert commentary from Jim Cowen.
###
Search, Discovery, and Answer Engines
WebSearches