Can AI grade AP free-response questions?

Updated 4 September 2026

For practice work, yes, as a first pass you review, and AP suits it better than most because AP rubrics are point rubrics rather than band rubrics: every point is earned or not earned against a written condition, so a first pass can name the condition and quote the sentence it judged. The mechanical points hold up, document count, evidence beyond the documents, contextualisation, a stated line of reasoning. The judgement points do not: a thesis point, a complexity or sophistication point, anything with the word nuanced in it. Those are the ones AP readers themselves argue about, and they are the ones to check by hand every time. No tool gives a student a score that counts. Only the College Board scores an AP exam.

Why point rubrics are different

Most writing rubrics are band rubrics: a piece is a 3 or a 4 on development, and the decision is a judgement about the whole. AP free response is not built that way. A response earns a thesis point, or it does not. It earns a contextualisation point, or it does not. Each point has a written condition attached.

That changes what a machine first pass is for. On a band rubric it produces a number you either accept or argue with. On a point rubric it produces a list of claims you can check one at a time, each attached to a sentence in the student’s own response. Checking six stated claims is much faster than forming one holistic judgement, and it is the reason a point rubric is the friendliest thing in this category to automate partially and the most dangerous thing to automate completely.

What holds up and what does not

AP free-response point types, how reliable an automated first pass is on each, and why
Point typeFirst passWhy
Document use and count (DBQ)ReliableIt is a counting question with a stated threshold
Evidence beyond the documentsReliablePresence of a specific outside example is checkable
ContextualisationMostly reliableNeeds a broader historical setting, stated rather than gestured at
Line of reasoningMostly reliableStructural, and visible in the response
Thesis or claimReview by handTurns on whether a claim is defensible and responsive, which is a judgement
Complexity or sophisticationReview by handThe point AP readers disagree about most; treat any machine verdict as a prompt to look
Commentary quality (English FRQ)Review by handThe difference between describing a device and explaining its effect is exactly the hard part

This table is our judgement about where an automated first pass is trustworthy, based on the shape of each point. It is not a measurement. We have no benchmark on AP-scored responses because the College Board does not release a scored set we could measure against, and we would rather say that than publish a number we cannot stand behind.

The method, five steps

  1. 1. Score against the published scoring guideline, not a paraphrase

    AP points are awarded against specific written conditions, and a paraphrased rubric quietly changes what counts. Load the real guideline for the exam and the question type you set.

  2. 2. Give it the prompt and the documents, not just the essay

    A DBQ cannot be scored without the documents, and a rhetorical analysis FRQ cannot be scored without the passage. A first pass judging evidence it has never seen is guessing, and it will guess confidently.

  3. 3. Take the mechanical points first

    Document count, evidence beyond the documents, contextualisation, a stated line of reasoning. These turn on whether something is present, they are fast to verify, and they are where most students actually lose marks.

  4. 4. Review every thesis and complexity point by hand

    These are the judgement points, the ones AP readers argue about, and the ones a first pass is least reliable on. Treat the machine score there as a prompt to look, never as a decision.

  5. 5. Give the student the condition, not the number

    A 4 out of 6 teaches nothing. "You did not earn the complexity point because you named a counterargument and never engaged with it" is the whole value of a point rubric, and it is the part worth spending your time on.

The AP and IB scoring guides we hold

Where this is the wrong tool

Anything that has to count. An AP score is awarded by the College Board. Nothing here is a score, a predicted score with any official standing, or a basis for an appeal.

Scoring without the source material. Hand it a DBQ response and no documents and it will produce a confident number that means nothing.

The judgement points, unsupervised. If you are not going to review thesis and complexity yourself, do not use a first pass for them at all. A wrong thesis point that nobody checks teaches a student the wrong lesson for a year.

Plagiarism and AI-writing detection. We do not do it. On take-home AP practice work that is a real gap and we have no answer to it.

No affiliation and no SOC 2. Not endorsed by the College Board or the International Baccalaureate. Institution plans are FERPA and COPPA aligned with a custom data processing agreement; there is no SOC 2 certification.

Common questions

Can AI grade AP free-response questions?

It can produce a point-by-point first pass that you review, and AP rubrics suit that better than most because they are point rubrics rather than band rubrics: each point is earned or not earned against a written condition, so a first pass can show you which condition it thinks was met and quote the sentence it based that on. What it cannot do is award a score that counts. Only the College Board scores an AP exam. Everything on this page is practice scoring for your own classroom.

Which points does it get right, and which does it get wrong?

The mechanical points hold up well: does the response describe a specific historical development, does it cite the required number of documents, does it use evidence beyond the documents, does it state a line of reasoning. Those turn on whether something is present, and presence is checkable. The judgement points are where it drifts. A thesis point on an LEQ, a sophistication or complexity point, a "nuanced" understanding: those are the ones AP readers themselves argue about at the reading table, and they are the ones to review first every single time.

Is there a measured accuracy figure for AP free response?

No, and we will not imply one. Our published benchmark is against 78 real student responses that Texas released with official examiner scores from the STAAR constructed-response scoring guides, where nine in ten came back within one point and quadratic weighted kappa was 0.78. That is a different exam, a different rubric shape and a different grade range. Treating it as an AP number would be dishonest. AP-scored responses with published official scores are not released as a set we can benchmark against.

Does it work on handwritten FRQ and DBQ responses?

Yes. Photograph the pages with a phone, one after another, and the response comes back marked in place. There is no template, no bubble sheet and no scanner. Anything genuinely unreadable is flagged for you rather than scored wrong, which matters more on a timed handwritten essay than almost anywhere else, because the handwriting is at its worst under time pressure.

Which AP rubrics do you hold?

The library holds the AP English Language FRQ rubric, the AP English Literature FRQ rubric, the AP History DBQ rubric, the AP History LEQ rubric, the AP Seminar Individual Written Argument and Individual Research Report scoring guidelines, and the AP Research academic paper guidelines, alongside 86 rubrics in total. They are reproduced from the published College Board scoring guidelines. You can also load your own version if your department scores differently in practice.

Is this endorsed by the College Board?

No. NudgeLearn has no affiliation with the College Board and this is not endorsed or approved by them. AP and Advanced Placement are their trademarks. We reproduce their published scoring guidelines so a teacher can score practice work against the real criteria, and nothing here produces an official score.

How should I use this with a class?

For practice sets and timed writes during the year, where the value is a student finding out in November which point they keep missing. Run five responses you have already scored yourself before you trust it with a set, and pay attention to whether it disagrees with you on the thesis point specifically, because that is the disagreement that will repeat.

What does it cost?

Free covers 20 graded items a month. Pro is $30 a month, or $25 billed yearly, for 400 items. Max is $50 a month, or $42 yearly, for 1,000. Essays and free responses bill by length: one item per 500 words, or one per page, whichever is larger. A class of 30 writing a 600-word LEQ is about 60 items. Extra items are 10 cents and overage is capped.

Try it on five you have already scored

Free for 20 graded items a month. Watch what it does with the thesis point in particular, because that is the disagreement that will repeat.

Grade a response free

Read next