Are AI Study Tools Actually Tested for Learning?
Accuracy Warning — AI study tools
Unguarded AI answer-giving can raise practice scores while lowering unassisted exam scores; no direct GRE/MCAT/SAT/ACT/ASVAB evidence.
- Accuracy:
- Moderate
- Tested:
- Practice with AI tutor followed by unassisted exam
- Last tested:
- 2026-08-03
Evidence label: mixed, by design. Last reviewed: August 3, 2026. Actual randomized trials of AI study tools now exist, so the useful question is no longer whether a product page can say “tested.” It is tested for what: higher practice accuracy while the bot is open, or better performance when the student has to solve alone?
For GRE, MCAT, ASVAB, SAT, and ACT prep, the honest answer is narrower than most marketing copy suggests. Unguarded general chatbot use has High evidence of risk in at least one field randomized trial. Guardrailed, purpose-built tutors have High to Moderate evidence of learning gains across several controlled settings. Vendor telemetry can be useful for product design, but it is Limited as proof that a student will score higher on a high-stakes exam.

| Claim type | StudyMethod evidence label | What the evidence actually tested | What exam preppers should take from it |
|---|---|---|---|
| Unguarded general chatbot for practice | High evidence of risk | In a field RCT of about 1,000 Turkish high school math students, GPT Base raised practice scores by 48% but reduced later unassisted exam scores by 17%; the most common prompt was “give me the answer.” [1] | A tool can make practice feel easier while weakening independent problem-solving. |
| Guardrailed AI tutor that scaffolds work | High to Moderate evidence of learning gains | The guardrailed GPT Tutor variant in the same PNAS study increased practice performance by 127% without the later exam harm reported for GPT Base; a separate Harvard crossover RCT found median learning gains more than double active learning in undergraduate physics. [1][2] | Design matters more than the brand name: the tutor has to make the student do the reasoning. |
| AI tutoring in an exam-prep-like setting | Moderate to High relevance, but not direct GRE/MCAT/SAT/ACT proof | An RCT with 334 university students preparing for an incentivized exam found a +0.23 SD gain versus textbook-only control. [3] | This is the closest match to prep behavior, but it still is not a high-stakes admissions or military placement test. |
| Vendor-announced or product telemetry results | Limited as outcome proof | Eedi/Google DeepMind, Google education pilots, and Khan Academy reports show promising design signals, but some are vendor-announced, pending further review, or based on product telemetry rather than independent exam outcomes. [4][5][6][7] | Useful clues for evaluating tool design; not enough for a score promise. |
The catch in “tested”: practice help is not the same as learning
The cleanest warning comes from Bastani and coauthors’ field experiment, because it separates the moment students feel helped from the moment they must perform without help. In that study, about 1,000 high school students in Turkey practiced math with different AI conditions and then took an unassisted exam. The GPT Base group did better during practice: practice scores rose 48%. Then came the part that matters for test prep: on the later exam, without the chatbot, their scores fell 17% relative to the control group. [1]
That is the result to keep in mind when an AI study tool says it improves performance. Performance where? If the metric is “got more practice questions right while the model was available,” the tool may be measuring assistance, not retained skill. For a student who will sit the GRE, MCAT, ASVAB, SAT, or ACT without the chatbot open, the final-assessment condition is not a technical detail. It is the point.
The mechanism was not mysterious. The most common student prompt in the GPT Base condition was “give me the answer,” and unguarded GPT-4 answered practice math problems correctly only 51% of the time. [1] That combination is almost designed to produce shallow fluency: the student sees something that looks like progress, may copy or follow the output, and loses the productive struggle that would have exposed gaps before exam day.

This is why “AI study tools tested for learning” cannot mean “students liked it,” “students completed more items,” or “the model benchmarked well.” Those may be useful signals for engagement or capability. They do not show that the learner can solve independently later.
The guardrailed version changed the learning result
The same PNAS experiment also tested a different design: GPT Tutor. That version was not just a friendlier chatbot skin. It was guardrailed to behave more like a tutor, pushing students through the problem rather than simply handing over the result. In that condition, practice performance increased by 127%, and the study did not find the same later exam harm seen in GPT Base. [1]
That contrast is more useful than a brand ranking. One design let students lean on the model as a crutch. The other constrained the interaction so the student still had to work. The result does not prove every guardrailed tutor works, and it does not prove a GPT Tutor-like system will raise your ACT math or GRE quant score. It does show that the answer-giving behavior itself is not a small usability quirk. It can change the learning outcome.

For exam prep, that pushes the buying question away from “Which model is smartest?” and toward “What can the tool refuse to do?” A useful tutor should be willing to slow the student down: ask for the next step, diagnose the error, reveal a hint before a solution, and make the learner commit to an answer. If the interface makes “show me the final answer” the fastest path, the student may still feel helped. The trial evidence says that feeling can be dangerous.
The strongest positive trials have scaffolding, verified work, or exam-prep structure
The Harvard crossover RCT by Kestin and coauthors is the positive result that deserves attention because it tested a purpose-built AI tutor against active learning in undergraduate physics. The study included 194 students and reported median learning gains more than double those from active learning, with effect sizes from 0.73 to 1.3 standard deviations. Students also rated the AI explanations highly: 83% said they were as good as or better than explanations from human instructors. [2]
The design details matter as much as the headline gain. The tutor used step-by-step scaffolding and pre-written verified solutions in the system prompt. [2] That is a different claim from “we attached a general chatbot to course content.” Verified worked solutions reduce the risk that the system teaches a polished wrong method. Scaffolding reduces the chance that students outsource the thinking.
The authors themselves warned against using AI as a crutch. [2] That warning fits the Bastani result almost too neatly: the same broad technology can help or hurt depending on whether it replaces the student’s reasoning or protects it.
The IZA exam-prep RCT is especially relevant for StudyMethod readers because its setting looks more like deliberate preparation than a one-off classroom novelty. In the trial summarized by Stanford SCALE, 334 university students prepared for an incentivized exam. AI tutoring produced a +0.23 SD gain versus textbook-only control, and unrestricted access beat restricted access by +0.21 SD. The summary also reports the largest gains for lower-baseline students with stronger self-regulation. [3]
That last detail should slow down anyone looking for a universal score promise. The students who benefited most were not simply “all students who touched AI.” They had lower starting performance and stronger self-regulation. [3] In test prep language, the tool may be most useful when a student can use feedback without letting it become answer delivery. A weak plan plus an always-available solver is not the same intervention.
Promising supporting evidence, with narrower labels
Eedi and Google DeepMind report an exploratory trial of a human-in-the-loop LearnLM tutoring system. In that announcement, the AI-supported approach matched expert tutors on mistake-fixing, 93.0% versus 91.2%, doubled knowledge-transfer gains, +10 percentage points versus +4.5 percentage points, and had a 0.1% factual-error rate. [4] Those are encouraging design signals: human oversight, attention to errors, and measurement beyond immediate correctness.
The label still cannot be High for exam-prep outcome proof. The results were announced by the organizations involved, and the broader evidence base is still developing. A second AI tutoring trial across 1,525 UK students was reported as beginning in 2026. [5] Until those results are fully available and reviewed, this is a promising, Moderate signal for AI tutor design, not a direct claim about SAT or MCAT score improvement.
Google’s Sierra Leone RCT broadens the pattern in a different direction. The reported study covered 48 classrooms and about 1,800 Grade 7–8 students, with a +0.26 SD gain on externally validated assessments and +0.38 SD past a 12-hour threshold. [6] That is stronger than ordinary engagement telemetry because it includes an assessment outcome. It is still Grade 7–8 math in Sierra Leone, not U.S. high-stakes admissions testing.
Khan Academy’s Khanmigo updates are useful in a different way. Khan reports that adding learning-history signals improved next-item correctness by 6.1%, and that response-conciseness design cut answer-giving by 50%. [7] Those numbers are worth reading as product-design telemetry: they suggest which interface choices may reduce answer dumping and improve immediate practice flow. They are not, by themselves, proof that a student will retain more on a proctored exam.
What none of these trials proves for GRE, MCAT, ASVAB, SAT, or ACT prep
None of the studies above ran under GRE, MCAT, ASVAB, SAT, or ACT conditions. Bastani was high school math in Turkey. Kestin was undergraduate physics at Harvard. Eedi was K–8 math in the UK. Sierra Leone was Grade 7–8 math. The IZA study is closest to exam prep because students were preparing for an incentivized exam, but it still is not one of the major high-stakes tests StudyMethod readers usually mean.
That boundary matters because transfer is the hard part. A tool that helps with algebra steps may not help with MCAT passage reasoning. A physics tutor that improves conceptual learning may not map cleanly onto GRE time pressure. A K–8 math intervention may say something about scaffolding but little about adult self-study habits over eight weeks. The fair conclusion is a design principle, not a named-tool guarantee.
If you are building an exam plan, the AI tool should sit inside an exam-first structure: diagnostic test, content review, timed sets, error log, full-length practice, and unassisted review. StudyMethod’s GRE prep by the numbers page is the kind of architecture to start from: time, cost, score movement, and practice-test behavior before tool enthusiasm.
A verification checklist for any “tested AI tutor” claim
When a study tool says it was tested for learning, ask for the study design before the feature list. The difference between a real learning claim and a soft marketing claim usually appears in five places.

- Was it randomized? A randomized controlled trial can support a causal claim better than usage data, testimonials, or before-and-after product metrics.
- Was the final assessment unassisted? If students used the AI during the outcome test, the study may show assisted performance rather than retained skill.
- Was the tool guardrailed against answer-giving? Look for step-by-step prompting, hint sequences, verified worked solutions, refusal to provide direct answers too early, and error diagnosis.
- Was there human oversight or verified content? Human-in-the-loop review and pre-verified solutions do not make a tool perfect, but they lower the risk of fluent wrong explanations.
- Who ran and reported the evidence? Independent peer-reviewed trials deserve more weight than vendor-run telemetry. Vendor data can still be useful, but it should not be treated as score-outcome proof.
- Did the outcome resemble your exam? A classroom concept quiz, a K–8 math assessment, and a proctored MCAT section are not interchangeable.
That checklist is also how to read model-specific claims. A benchmark can show that a model solves many tasks; it does not automatically show that a student learns more after using it. For examples of how StudyMethod separates benchmark strength from exam-prep usefulness, see the evidence-labeled reviews of Claude for exam prep, Claude study safety, and DeepSeek benchmark transfer.
How to use AI without letting it become the hidden test-taker
The safest use pattern is simple: let AI tutor the process, then remove it before the score matters. Ask for a hint, a diagnosis, a simpler analogous problem, or a check of your reasoning. Do not let it become the fastest route to the final answer on every missed question.
For a missed practice item, a guardrailed process might look like this: first attempt the problem alone; ask the AI to identify the concept being tested without solving it; write your next step; ask for feedback on that step; finish the item; then compare your work to an official or verified solution. After that, put the mistake into an error log and retest the concept without AI later. The unassisted retest is where learning gets checked.
There are still legitimate uses for general chatbots in prep: generating a study schedule, explaining a concept at a different reading level, creating a short set of untimed drills, or helping organize an error log. The boundary is answer substitution. StudyMethod’s guide to whether students can use ChatGPT for exam prep draws that safe-use line more directly, and the Gen Beta AI learning tools piece applies the same verification-first rule to tool selection.
The current trial evidence justifies choosing tools that scaffold, verify, and withhold answers long enough for the student to think. It does not justify a promise that any named AI study tool will raise a GRE, MCAT, ASVAB, SAT, or ACT score. The student, not the chatbot, still has to sit for the exam.
References
- Generative AI can harm learning, PNAS, June 2025
- AI tutoring outperforms active learning, Scientific Reports, June 2025
- AI Tutoring Enhances Student Learning Without Crowding Out Reading Effort, Stanford SCALE
- New exploratory research from Eedi and Google DeepMind reveals human-in-the-loop AI tutoring outperforms human-only support, Eedi
- Eedi and Google DeepMind begin second AI tutoring trial across 1,525 UK students, EdTech Innovation Hub, 2026
- Measuring the impact of AI on teaching and learning, Google
- How Khan Academy is building a better AI tutor: Our most recent learnings, Khan Academy
Authoritative source
For the authoritative version of this contentHow to Read the '1 in 4 NFL Players CTE' Study
Report an error in this tool's output
Found something this tool got wrong beyond what's documented above? Report it so the accuracy log stays current.

Comments
Join the discussion with an anonymous comment.