AmberPen Get started

Benchmarks & performance

We evaluate AmberPen on public academic benchmarks and run every competitor against the same corpora with the same scorers. Two datasets cover complementary failure modes: JFLEG measures fluency corrections (GLEU, higher is better) and CWEB measures restraint on low-error-density web text (ERRANT F0.5 — precision and recall over individual edits).

Correction quality — JFLEG

754 sentences from the JFLEG development set, scored with the official GLEU implementation against all four human references. Each row is a complete run.

ProofreaderGLEU95% interval
AmberPen0.5610.544–0.578
Sapling0.4700.454–0.486
GrammarBot Neural0.4130.400–0.425
Harper 2.4.00.3440.328–0.359
LanguageTool0.3070.293–0.322

Precision on clean text — CWEB

6,845 web sentences where most text is already correct, so over-correcting is punished. ERRANT F0.5 weighs precision twice as heavily as recall.

ProofreaderPrecisionRecallERRANT F0.5
GrammarBot Neural40.8%62.5%43.9
AmberPen37.4%65.5%40.9
LanguageTool37.8%18.3%31.2
Sapling23.6%49.1%26.3
Harper 2.4.05.8%13.8%6.6

GrammarBot Neural leads CWEB F0.5 by 2.99 points. AmberPen has the highest recall, finding 3.6× more of the annotated errors than LanguageTool at similar precision, and remains the strongest product on JFLEG fluency.

Speed

End-to-end latency of each corrector call during the complete benchmark runs, measured per sentence:

Proofreaderp50p95p99
Harper 2.4.0 (local)2 ms8 ms14 ms
AmberPen217 ms1,024 ms1,413 ms
LanguageTool340 ms504 ms602 ms
Sapling620 ms1,086 ms1,646 ms
GrammarBot Neural1,000 ms2,408 ms4,421 ms

Local Harper holds the lowest raw latency. Its in-process timing is not directly comparable to a hosted API call; among hosted correctors, LanguageTool is the most regular. AmberPen is roughly 3× faster than Sapling and 4.6× faster than GrammarBot Neural at the median. Latencies are from the CWEB run (single sentences, ~100 characters); longer inputs take proportionally longer.

Quality, price, and speed

The combined quality score is the arithmetic mean of JFLEG GLEU × 100 and CWEB ERRANT F0.5, giving the two benchmarks equal weight on a 0–100 scale. Higher is better. The shaded upper-left area marks the most attractive quadrant: higher quality at lower cost or latency.

Quality vs price

Most attractive quadrant
Quality vs priceA scatter plot where higher quality and positions farther left are more attractive.20304050$0$10$20$30Harper: 20.5 quality, $0HarperAmberPen: 48.5 quality, $9AmberPenSapling: 36.7 quality, $25SaplingGrammarBot Neural: 42.6 quality, $30GrammarBot NeuralCost per 1M characters (USD)Combined quality score

Entry paid-plan rates. Harper is free and local. LanguageTool is omitted because it is priced per request rather than per character.

Quality vs speed

Most attractive quadrant
Quality vs speedA scatter plot where higher quality and positions farther left are more attractive.203040500 ms300 ms600 ms900 ms1200 msHarper: 20.5 quality, 2 msHarperAmberPen: 48.5 quality, 217 msAmberPenLanguageTool: 31.0 quality, 340 msLanguageToolSapling: 36.7 quality, 620 msSaplingGrammarBot Neural: 42.6 quality, 1000 msGrammarBot NeuralMedian latency per sentenceCombined quality score

Median end-to-end latency from the CWEB run. Harper ran in-process; all other products used hosted APIs.

Method

  • All correctors received identical inputs; Harper ran locally and the others used their APIs.
  • JFLEG scored with the official GLEU scorer and all four references.
  • CWEB scored with ERRANT 3.0.2 (spaCy 3.8); gold edits regenerated from CWEB's raw human references with the same tokenizer as the hypotheses.
  • AmberPen ran in correct mode at temperature 0.
  • Latency percentiles are computed over individual corrector calls, excluding benchmark-runner queue time. Complete runs, July 2026.

Comparisons