Do you need a scanner to grade paper exams with AI?

Updated 4 September 2026

Not if you are one teacher grading one class. A phone photo of each page is enough, and a whole set can be photographed one page after another. There is no template to print, no bubble sheet, no fixed answer box and no code on the paper, because a photo-based grader reads an exam that was written for students rather than for a machine. The tools that do need a scanner and a printed form are scanning platforms, and they need it for a reason: they are built to grade one exam across an entire cohort with roster matching, and at that scale the template is the feature rather than the cost. If you are grading hundreds of identical forms across a department, use one of those. If you are grading thirty papers tonight, the scanner is apparatus you do not need.

What the apparatus actually costs you

The hidden price of a template-based workflow is not the scanner. It is that the exam has to be designed for the tool before a student ever sits it: answers in fixed boxes, a code on every page, a reprint if the layout changes, and a quiz you wrote at 8am on the back of an envelope that cannot go through the system at all.

That price is worth paying at cohort scale, where a department grades one common exam and consistency across hundreds of scripts is the whole point. It is not worth paying for a Friday test in one classroom, which is the case nothing in the category is built for and the reason this product exists.

The method, six steps

  1. 1. Do not redesign the test

    You do not need to move the answers into boxes, add a QR code, or reprint anything. The point of a photo-based grader is that it reads a paper that was written for students rather than for a machine. If a tool asks you to change the exam first, that tool is a scanning platform and it is a different product.

  2. 2. Photograph every page in handover order

    Phone camera, flat page, whole page in frame, decent light. Multi-page exams stay in order per student. Do not stop to mark anything; capturing the whole set in one continuous pass is what stops thirty exams from becoming thirty separate jobs.

  3. 3. Let it work each question rather than compare to a key

    On an exam this decides how partial credit lands. Solving each question independently means an equivalent form still counts, units are handled, and a wrong step under a right final answer is caught. A string-matching key marks the first as wrong and the second as right, and both mistakes cost a student marks they earned or hide a misconception they have.

  4. 4. Deal with the flagged answers first

    Anything genuinely unreadable comes back flagged rather than scored. On an exam this matters more than on homework because the handwriting is at its worst under time pressure and the stakes are highest. A grader that never flags anything is guessing, and a quietly guessed answer marks a correct student wrong with nothing looking broken.

  5. 5. Set partial credit yourself where it matters

    Machine partial credit is a starting point, not a policy. Your department has a convention about how much a sign error costs, and no tool knows it. Override the ones that matter, and only the ones that matter: the boundary answers are the only ones a student will ever query.

  6. 6. Release, then read the per-question data

    Nothing reaches a student until you release it. Once it has, the per-question breakdown by student and class is what tells you which question the exam actually tested badly, which is the one piece of exam data most teachers never get because assembling it by hand is not worth the hours.

Partial credit without an answer key

An exam lives or dies on partial credit, and an answer key cannot award it. A key compares the final answer to a string: a student who wrote 0.5 where the key says one half loses marks they earned, and a student who lost a sign in step two and recovered by accident keeps marks they did not.

Working each question out independently gives you the thing a key cannot: the final answer is right and line three is wrong, stated in plain English, on the student’s own page. That is a defensible starting point for partial credit. It is a starting point. Your department has a convention about what a sign error costs and no tool knows it, so every allocation is editable and nothing reaches a student until you release it.

Where a scanning platform genuinely beats this

One common exam across a whole cohort. Hundreds of identical scripts, multiple graders, roster matching, question-by-question grading across the cohort at once. A platform built for that will beat us at it, and we say so on our own comparison page rather than waiting for you to find out.

Pure multiple-choice at volume. If the exam is all bubbles and there are several hundred, a bubble-sheet scanner is the right instrument. Photo grading earns its keep on the written working, which a bubble sheet cannot capture at all.

Anything your district has already procured. We hold no SOC 2 certification. Institution plans are FERPA and COPPA aligned with a custom data processing agreement, and if SOC 2 is a requirement that settles it before any feature comparison starts.

The full head-to-head, including their free tier

Where this goes wrong

Bad photographs. Faint pencil, heavy crossing-out and pages shot at an angle all cost accuracy. Flat page, decent light, whole page in frame.

Freehand diagrams, graphs and constructions. Read far less reliably than working and text. On a geometry exam, expect to check those yourself.

No published math benchmark. Our measured accuracy is on writing, not on math, and an exam is the worst place to assume a number transfers. Run five you have already graded first.

No plagiarism or AI-writing detection. We do not do it, and for a take-home exam that is a real gap we have no answer to.

Common questions

Do you need a scanner to grade exams with AI?

Not with a photo-based grader. A phone photo of each page is enough, and a whole class set can be photographed one page after another. There is no template to print, no bubble sheet, no fixed answer box and no QR code on the paper. The tools that do require a scanner and a printed template are scanning platforms, which are built for grading one exam across a whole cohort rather than one class, and for that job the template is the feature rather than the cost.

Can AI grade a handwritten paper test?

Yes. It reads pencil, pen, cramped working, crossings-out and answers written in the margin, and it reads multi-step working rather than just a boxed final answer. Marks land on the student’s own page with a short reason for anything wrong. Anything genuinely unreadable is flagged for you to confirm rather than scored, which is the behaviour to test on any tool you are considering, ours included.

How does partial credit work with no answer key?

Each question is worked out independently rather than matched against a key, so it can distinguish a right answer reached by a wrong route from a wrong answer reached by sound reasoning with one arithmetic slip. That gives you a defensible starting point for partial credit. It is a starting point: your department has conventions about what a sign error costs and no tool knows them, so the machine allocation is fully editable and nothing goes to a student until you release it.

What about multiple choice sections?

Those are read from the page like anything else, with no bubble sheet required. If your exam is entirely multiple choice and you are grading several hundred across a department, a bubble-sheet scanner is genuinely the better instrument and we will not pretend otherwise. Photo grading earns its keep on the written working, which is the part a bubble sheet cannot capture at all.

When is a scanning platform the better choice?

When every student answers in the same box on the same printed form and you are grading at cohort scale across a department, with roster matching and question-by-question grading over hundreds of scripts at once. That is a real job and a platform built for it will beat us at it. Gradescope publishes no price for that tier; Gradescope Basic is free and their AI-powered grading sits in the quoted institutional plan. Our comparison page says the same thing.

Is it accurate enough for a summative exam?

Treat it as a first pass you review, on a summative exam more than anywhere else. Our published accuracy figure is on writing, not on math: 78 real student responses that Texas released with official examiner scores, nine in ten within one point of the examiner and quadratic weighted kappa 0.78 against 0.70 as the commonly accepted bar. There is no equivalent published benchmark for math yet and we say so rather than implying one. Run five exams you have already graded before you trust it with a set that counts.

What does it cost to grade a class of thirty?

Billing is by graded item, and a page is an item, so a one-page test for thirty students is thirty items and a four-page exam for thirty is 120. Free covers 20 items a month. Pro is $30 a month, or $25 billed yearly, for 400. Max is $50 a month, or $42 yearly, for 1,000. Extra items are 10 cents each and overage is capped, so an exam week cannot become a surprise bill.

Test it on the exam you already graded

A free account grades 20 items a month. Take five scripts you have already marked, including the two you argued with yourself about, and compare.

Start grading free

Read next