Back to articles
Everyone Got the Same AI and the Gap Did Not Move

Everyone Got the Same AI and the Gap Did Not Move

AI access is now near-universal, yet outcomes still split. The divide moved from who has the tool to who keeps doing the work.

October 2, 2026 · 9 min read
Add as a preferred source on Google

If you run a platform or own a training budget, you probably carry one assumption into every AI rollout: give everyone the same tool and the playing field levels.

Indonesia is a clean test of that assumption. There, 53% of 15-year-olds use an AI chatbot to help them learn at least once a week, against an OECD average of 46%. Only 7% say they never use one for assigned work, half the OECD figure of 14%.

That looks like a head start. The same OECD country note, released with the PISA 2025 results on 8 September 2026, reports that 37% of Indonesian 15-year-olds reach Level 2 or higher in science. The OECD average is 74%.

I read the evidence from five countries the same way. Access has gone flat, and the gap now sits in a place your adoption dashboard does not look.

Key takeaways

  • Near-universal access to AI has not closed outcome gaps. In Indonesia, 53% of 15-year-olds use AI weekly to learn, above the OECD average of 46%, while 37% reach Level 2 in science against 74% across the OECD.
  • AI usage intensity does not predict performance. A Dutch study of 4,497 users found no significant link between AI usage and test results, a clear link for digital literacy, and a family-background advantage that persisted regardless of AI.
  • Output scores stop measuring learning once AI arrives. In a Chinese panel of 26,811 users, homework scores rose 18% and closed-book exam scores fell 20% within six months of adopting generative AI.
  • The effort floor check separates assisted work from outsourced work. Compare AI users’ time on task with the fastest unaided users. Work finished below that floor with higher scores is outsourcing, and those scores should not count as progress.
  • Supervision and disposition decide who benefits from AI, and both are unevenly distributed. A product that leaves effort optional hands the advantage to users who already have structure around them.

Hierarchy showing same AI, unequal outcomes, branching into access, disposition, supervision and effort.

Does equal access to AI level outcomes?

No. Across the sources behind this piece, access is the one variable that no longer varies, and outcomes still do.

Start with the newest evidence. On 25 August 2026, Chalkbeat reported on a two-year experiment across 18 sites in one Tennessee district, where low-performing users worked in Khan Academy with the Khanmigo AI assistant one click away. They opened it on about a third of the days they were on the platform. When they did, many sent off-topic messages or tried to get it to hand over the answer. “Access was nearly universal but engagement was thin,” the researchers wrote. The platform produced faster math gains than the comparison group by year two, and the researchers say that benefit did not seem to come from the AI.

A 2026 Dutch study by Wang and colleagues makes the same point with registry data. It links surveys to national tax and test records for 4,497 users in their final primary year, in a system where adaptive AI tools are standard. AI usage intensity had no significant relationship with test performance (B = 0.03). Digital literacy did (B = 1.78). Family socio-economic status kept a strong direct link to performance (B = 0.23) that ran through neither. The tool reached everyone, and the oldest gap in the data stayed where it was.

Where did the divide go?

The divide moved from having the tool to what users do with it, and that is sorted by disposition and by supervision.

The Dutch data shows the disposition half. Openness, extraversion and perseverance each predicted digital literacy (B = 0.19, 0.17 and 0.05), and digital literacy carried part of their link to performance. Family background barely predicted digital literacy at all; the coefficient was slightly negative. The authors add a caveat: the data dates from spring 2023, before generative AI was widely integrated, and it captures institution-sanctioned tools only.

The supervision half shows up in what users do on their own. An Alan Turing Institute survey of 780 UK children aged 8 to 12 found that 52% of those at private institutions had used generative AI, against 18% at state-funded ones. Among staff aware of AI use in assigned work, 47% at private institutions reported users submitting AI-generated work as their own, against 60% at state-funded ones. More use, and by that report less substitution, where there is more adult structure around the user. These are self-reports from one country, so I treat them as a direction only.

AI skill has not become a marker of family income, at least in the one study that measured it directly. What is unequal is the structure around the user: who watches the inputs, and who keeps working when the answer is one prompt away. Structure costs money and attention, which is why I expect it to track advantage over time, the same way access never guaranteed participation in emerging markets.

Decision flow for the effort floor check: find the unaided time floor, count AI users below it, then decide which score to trust.

What does outsourcing cost when nobody is watching?

About a fifth of measured learning within six months, in a large natural-setting study.

Strömberg, Lei and Wu tracked 26,811 secondary-level users in one Chinese county for 30 months, as reported AI use went from nearly zero to around 80%. In their June 2026 working paper, adoption raised homework scores by 18% and cut completion time from 64 to 45 minutes. Scores on monthly closed-book exams fell by 20% within six months. High-stakes entrance exam scores fell by 18% and 24%, and that loss took about two years to show in full.

The paper reports adopters and non-adopters as well balanced on demographics and prior exam scores. The losses concentrated in the roughly 80% of AI users whose behaviour was consistent with outsourcing. AI users who kept spending as much time on homework as non-users reached similar exam scores, and they were not stronger performers to begin with.

The paper also undercuts the comfortable version of my argument. Losses were larger for high achievers: 24% for the top third by prior score, against 16% for the bottom third. Unsupervised AI pulled the strongest users down furthest. That is why I read the divide as one of effort and oversight first, and of income only where income buys oversight.

One finding belongs in every product review. Among AI users with above-median homework scores, higher homework scores went with lower exam scores. The output metric had inverted, which is the sharpest case yet for designing the struggle back into a learning platform.

Bar chart: Indonesia leads the OECD average on weekly AI use, 53 to 46 percent, and trails on Level 2 science, 37 to 74 percent.

How do you tell assisted work from outsourced work?

Run what I call the effort floor check. It comes straight from how the Chinese data separates the two kinds of AI user.

  1. Find the floor. Take the fastest time in which unaided users complete the task. In the study, non-AI users generally needed at least 50 minutes per assignment.
  2. Count who is under it. More than half of AI users finished in 20 to 50 minutes, faster than even the fastest unaided users. The authors classed anything under 50 minutes as outsourcing: 58% of AI users overall, and 81% after more than five months of use.
  3. Compare the two scores. Under the floor, homework scores were very high and exam scores extremely low. Between 50 and 65 minutes, AI and non-AI users had similar exam scores.

The decision rule: if a majority of AI-assisted users finish below the unaided floor while their output scores rise, stop counting output scores as progress and move the measurement to unaided performance.

The check travels beyond formal learning. Any capability programme has an output metric and a later unaided test. Onboarding is the obvious place to start, because the juniors you automate are the seniors you cannot hire.

The paper’s own chart of scores against completion time shows where that floor sits.

Two charts: AI users finishing homework in under 50 minutes score about 120 on homework and 70 to 80 on exams.

Exhibit 1. Homework and exam scores by homework completion time, for users of generative AI and for non-users. The solid red line is users who had adopted generative AI, the dashed blue line is users who had not, and the shaded bands are interquartile ranges. Scores are a percentage of the pre-AI baseline mean, so 100 is the average non-AI user. Between about 25 and 45 minutes the red line sits near 120 in the top panel (Homework Score) and mostly between 70 and 80 in the bottom panel (Exam Score); the dashed line only begins at about 50 minutes, near 100 in the top panel and in the 90s in the bottom panel. Source: Strömberg, D., Lei, V., & Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. Working paper, June 2026, Figure 5, p. 34, p. 34.

Flow from flat access to what users do with the tool, split into disposition and supervision, ending in outcomes that still differ.

What should a product team change?

Design for the user who will not volunteer the effort, because that is most users.

Since the study period, Khan Academy has redesigned its interface to integrate the assistant, and Sal Khan wrote that the team “had to make productive struggle harder to sidestep.” The Chinese authors reach the same place from the demand side: AI tools built to coach already exist there at low or zero cost, and most users pick general-purpose tools that give quick answers. They suggest monitoring inputs such as homework time over outputs such as homework scores.

For a product team, I’d turn that into four changes:

  • Effort as the default path. Put the assistant inside the task, with steps that cannot be skipped, so the user without supervision gets the same structure as the user with it.
  • Input telemetry next to output scores. Report time on task against the unaided floor for each cohort, so an inverted metric shows up in weeks.
  • Unaided checkpoints on a schedule. The entrance-exam loss took two years to surface in full. A regular closed-book check is the cheapest early warning you can buy.
  • Skill-building budgeted apart from access. The Dutch result ties performance to digital literacy and finds nothing for usage, which is the case for treating critical thinking with AI as a separate, measurable skill.

The Indonesian numbers are the reason to act now. A country can sit above the OECD average on AI use and at half the OECD share on baseline science proficiency in the same survey. The OECD note does not claim one causes the other, and neither do I. If you own an AI rollout, run the effort floor check on one cohort this quarter before you report adoption as a win.

References

Frequently asked questions

Does giving everyone access to AI tools reduce outcome gaps?

No. A 2026 Dutch study of 4,497 users found AI usage intensity had no significant relationship with test performance, while family socio-economic status kept a strong direct link to performance regardless of AI. Access is now near-universal and outcomes still differ.

What is the effort floor check?

The effort floor check compares AI users' time on task with the fastest time unaided users need. Work completed below that floor with higher output scores is treated as outsourcing, so its scores should not be counted as progress. It is David Yosuanto's synthesis of a 2026 working paper by Strömberg, Lei and Wu.

How much does outsourcing work to generative AI cost in learning?

In a 30-month panel of 26,811 secondary-level users in China, adopting generative AI raised homework scores by 18% and lowered monthly closed-book exam scores by 20% within six months. Users who kept spending as much time on homework as non-users reached similar exam scores.

Is AI literacy determined by family income?

Not directly, in the evidence available. The Dutch study found family socio-economic status had a negligible relationship with digital literacy, while openness, extraversion and perseverance predicted it. Income matters where it buys supervision and structure around the user.

Evidence

They opened it on about a third of the days they were on the platform.

Students only used Khanmigo about a third of the days they were working in Khan Academy

Barnum, M. (2026, August 25). The problem with an AI tutor: Students don't want to use it. Chalkbeat. (p. 3)

The platform produced faster math gains than the comparison group by year two, and the researchers say that benefit did not seem to come from the AI.

Khan Academy made faster math gains than the comparison group … seem to come from AI

Barnum, M. (2026, August 25). The problem with an AI tutor: Students don't want to use it. Chalkbeat. (p. 3)

Under the floor, homework scores were very high and exam scores extremely low.

This group receives very high homework scores … they receive extremely low exam scores

Strömberg, D., Lei, V., & Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. Working paper, June 2026. (p. 18)

An Alan Turing Institute survey of 780 UK children aged 8 to 12 found that 52% of those at private institutions had used generative AI, against 18% at state-funded ones.

52% of children attending private schools report using generative AI, as opposed to 18% of children in state schools

The Alan Turing Institute. Understanding the Impacts of Generative AI Use on Children. Project summary. (p. 3)

It links surveys to national tax and test records for 4,497 users in their final primary year, in a system where adaptive AI tools are standard.

final analytic sample consisted of 4497 … SES was measured by a composite index derived from national … tax registry data, combining 50% parental-average wealth percentile … end-of-primary-education exit test scores from national registry

Wang, Z., van Wetten, S., Segers, E., & Haelermans, C. (2026). Decoding divides: The role of socioeconomic status and personality traits in AI divides and educational inequality. Computers and Education: Artificial Intelligence, 10, 100566. (p. 4)

Family socio-economic status kept a strong direct link to performance (B = 0.23) that ran through neither.

rect correlation with academic performance in the mediation model … = 0.23,

Wang, Z., van Wetten, S., Segers, E., & Haelermans, C. (2026). Decoding divides: The role of socioeconomic status and personality traits in AI divides and educational inequality. Computers and Education: Artificial Intelligence, 10, 100566. (p. 6)

Family background barely predicted digital literacy at all; the coefficient was slightly negative.

tistically significant yet close-to-zero negative relationship with digital

Wang, Z., van Wetten, S., Segers, E., & Haelermans, C. (2026). Decoding divides: The role of socioeconomic status and personality traits in AI divides and educational inequality. Computers and Education: Artificial Intelligence, 10, 100566. (p. 6)

When they did, many sent off-topic messages or tried to get it to hand over the answer.

many students just sent off-topic messages or tried to convince Khanmigo to provide the correct answers

Barnum, M. (2026, August 25). The problem with an AI tutor: Students don't want to use it. Chalkbeat. (p. 3)

The authors classed anything under 50 minutes as outsourcing: 58% of AI users overall, and 81% after more than five months of use.

student as engaging in homework outsourcing if the student completes homework in less than 50 minutes … 58 percent of AI students engage in homework outsourcing … After more than …ve months of generative AI use, these shares rise to 81 percent

Strömberg, D., Lei, V., & Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. Working paper, June 2026. (p. 18)

The paper reports adopters and non-adopters as well balanced on demographics and prior exam scores.

ever-adopters and never-adopters are well balanced on all the demographic variables as well as students

Strömberg, D., Lei, V., & Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. Working paper, June 2026. (p. 9)

In the study, non-AI users generally needed at least 50 minutes per assignment.

For non-AI students, homework completion generally takes at least 50 minutes

Strömberg, D., Lei, V., & Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. Working paper, June 2026. (p. 17)

Scores on monthly closed-book exams fell by 20% within six months.

scores in monthly school exams fall by 20 percent of the baseline mean

Strömberg, D., Lei, V., & Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. Working paper, June 2026. (p. 3)

The authors add a caveat: the data dates from spring 2023, before generative AI was widely integrated, and it captures institution-sanctioned tools only.

2023, just before generative AI became widely integrated … institutionally sanctioned

Wang, Z., van Wetten, S., Segers, E., & Haelermans, C. (2026). Decoding divides: The role of socioeconomic status and personality traits in AI divides and educational inequality. Computers and Education: Artificial Intelligence, 10, 100566. (p. 9)

Among AI users with above-median homework scores, higher homework scores went with lower exam scores.

Among AI students with above-median homework scores, higher homework scores are associated with lower exam scores

Strömberg, D., Lei, V., & Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. Working paper, June 2026. (p. 20)

On 25 August 2026, Chalkbeat reported on a two-year experiment across 18 sites in one Tennessee district, where low-performing users worked in Khan Academy with the Khanmigo AI assistant one click away.

worked with a Tennessee school district to randomly assign low-performing students from 18 middle schools to use Khan Academy

Barnum, M. (2026, August 25). The problem with an AI tutor: Students don't want to use it. Chalkbeat. (p. 3)

Among staff aware of AI use in assigned work, 47% at private institutions reported users submitting AI-generated work as their own, against 60% at state-funded ones.

47% of teachers working in private schools report awareness around this type of use by their students, compared to 60% of teachers in state schools

The Alan Turing Institute. Understanding the Impacts of Generative AI Use on Children. Project summary. (p. 4)

AI usage intensity had no significant relationship with test performance (B = 0.03).

sity showed no significant relationship with academic performance … 0.03,

Wang, Z., van Wetten, S., Segers, E., & Haelermans, C. (2026). Decoding divides: The role of socioeconomic status and personality traits in AI divides and educational inequality. Computers and Education: Artificial Intelligence, 10, 100566. (p. 6)

The Chinese authors reach the same place from the demand side: AI tools built to coach already exist there at low or zero cost, and most users pick general-purpose tools that give quick answers.

such tutoring tools already exist and are often available at low or zero cost … rely on general-purpose AI tools that provide quick answers

Strömberg, D., Lei, V., & Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. Working paper, June 2026. (p. 23)

Losses were larger for high achievers: 24% for the top third by prior score, against 16% for the bottom third.

(-24 percent) for the highest tercile and the least negative (-16 percent) for the lowest tercile

Strömberg, D., Lei, V., & Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. Working paper, June 2026. (p. 22)

Digital literacy did (B = 1.78).

a significantly positive relationship with academic performance … 1.78,

Wang, Z., van Wetten, S., Segers, E., & Haelermans, C. (2026). Decoding divides: The role of socioeconomic status and personality traits in AI divides and educational inequality. Computers and Education: Artificial Intelligence, 10, 100566. (p. 6)

"Access was nearly universal but engagement was thin," the researchers wrote.

universal but engagement was thin

Barnum, M. (2026, August 25). The problem with an AI tutor: Students don't want to use it. Chalkbeat. (p. 3)

AI skill has not become a marker of family income, at least in the one study that measured it directly.

acy relationship was statistically significant but practically negligible

Wang, Z., van Wetten, S., Segers, E., & Haelermans, C. (2026). Decoding divides: The role of socioeconomic status and personality traits in AI divides and educational inequality. Computers and Education: Artificial Intelligence, 10, 100566. (p. 7)

AI users who kept spending as much time on homework as non-users reached similar exam scores, and they were not stronger performers to begin with.

AI users who spend as much time on homework as non-users achieve similar exam scores … selected on prior achievement

Strömberg, D., Lei, V., & Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. Working paper, June 2026. (p. 3)

They suggest monitoring inputs such as homework time over outputs such as homework scores.

monitor inputs, such as homework time and study … rather than outputs such as homework scores

Strömberg, D., Lei, V., & Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. Working paper, June 2026. (p. 23)

There, 53% of 15-year-olds use an AI chatbot to help them learn at least once a week, against an OECD average of 46%.

53% of students reported using AI chatbots to help them learn at least once a week (OECD average: 46%)

OECD (2026). PISA 2025 Results (Volume I): Indonesia. Country note, 8 September 2026. (p. 12)

In a Chinese panel of 26,811 users, homework scores rose 18% and closed-book exam scores fell 20% within six months of adopting generative AI.

30 months of panel data on 26,811 Chinese students … raises homework scores by 18% … monthly exam scores by 20% within six months

Strömberg, D., Lei, V., & Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. Working paper, June 2026. (p. 1)

More than half of AI users finished in 20 to 50 minutes, faster than even the fastest unaided users.

More than half of AI students spend 20 … less time than even the fastest non-AI students

Strömberg, D., Lei, V., & Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. Working paper, June 2026. (p. 18)

The losses concentrated in the roughly 80% of AI users whose behaviour was consistent with outsourcing.

learning losses are concentrated among the roughly 80 percent of AI students who spend substantially less time on homework than non-AI students

Strömberg, D., Lei, V., & Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. Working paper, June 2026. (p. 19)

Openness, extraversion and perseverance each predicted digital literacy (B = 0.19, 0.17 and 0.05), and digital literacy carried part of their link to performance.

positive and substantively larger (openness: 𝐵 = 0.19, extraversion: … 𝐵 = 0.17, perseverance: 𝐵 = 0.05) … digital literacy significantly yet partially mediated the

Wang, Z., van Wetten, S., Segers, E., & Haelermans, C. (2026). Decoding divides: The role of socioeconomic status and personality traits in AI divides and educational inequality. Computers and Education: Artificial Intelligence, 10, 100566. (p. 7)

In their June 2026 working paper, adoption raised homework scores by 18% and cut completion time from 64 to 45 minutes.

homework scores rise by 18 percent of the baseline … completion time per assignment falls from 64 to 45 minutes

Strömberg, D., Lei, V., & Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. Working paper, June 2026. (p. 3)

Since the study period, Khan Academy has redesigned its interface to integrate the assistant, and Sal Khan wrote that the team "had to make productive struggle harder to sidestep."

Khan Academy has redesigned its interface to better integrate … We had to make productive struggle harder to sidestep

Barnum, M. (2026, August 25). The problem with an AI tutor: Students don't want to use it. Chalkbeat. (p. 4)

A Dutch study of 4,497 users found no significant link between AI usage and test results, a clear link for digital literacy, and a family-background advantage that persisted regardless of AI.

4497 Grade 6 students in the Netherlands … SES advantages persisted regardless of classroom AI implementation we observed

Wang, Z., van Wetten, S., Segers, E., & Haelermans, C. (2026). Decoding divides: The role of socioeconomic status and personality traits in AI divides and educational inequality. Computers and Education: Artificial Intelligence, 10, 100566. (p. 1)

Only 7% say they never use one for assigned work, half the OECD figure of 14%.

Only 7% of students reported never or almost never using AI chatbots for any of the schoolwork tasks examined in PISA (OECD average: 14%)

OECD (2026). PISA 2025 Results (Volume I): Indonesia. Country note, 8 September 2026. (p. 12)

Strömberg, Lei and Wu tracked 26,811 secondary-level users in one Chinese county for 30 months, as reported AI use went from nearly zero to around 80%.

30 months of panel data on 26,811 Chinese students

Strömberg, D., Lei, V., & Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. Working paper, June 2026. (p. 1)

The same OECD country note, released with the PISA 2025 results on 8 September 2026, reports that 37% of Indonesian 15-year-olds reach Level 2 or higher in science.

Some 37% of students in Indonesia attained Level 2 or higher in science (OECD average: 74%)

OECD (2026). PISA 2025 Results (Volume I): Indonesia. Country note, 8 September 2026. (p. 7)

High-stakes entrance exam scores fell by 18% and 24%, and that loss took about two years to show in full.

it takes two years for the negative … of 18 and 24 percent of the baseline mean

Strömberg, D., Lei, V., & Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. Working paper, June 2026. (p. 3)

Between 50 and 65 minutes, AI and non-AI users had similar exam scores.

of exam scores of AI and non-AI students are similar

Strömberg, D., Lei, V., & Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. Working paper, June 2026. (p. 18)