Benchmarks & performance
We evaluate AmberPen on public academic benchmarks and run every competitor against the same corpora with the same scorers. Two datasets cover complementary failure modes: JFLEG measures fluency corrections (GLEU, higher is better) and CWEB measures restraint on low-error-density web text (ERRANT F0.5 — precision and recall over individual edits).
Correction quality — JFLEG
754 sentences from the JFLEG development set, scored with the official GLEU implementation against all four human references. Each row is a complete run.
| Proofreader | GLEU | 95% interval |
|---|---|---|
| AmberPen | 0.561 | 0.544–0.578 |
| Sapling | 0.470 | 0.454–0.486 |
| GrammarBot Neural | 0.413 | 0.400–0.425 |
| Harper 2.4.0 | 0.344 | 0.328–0.359 |
| LanguageTool | 0.307 | 0.293–0.322 |
Precision on clean text — CWEB
6,845 web sentences where most text is already correct, so over-correcting is punished. ERRANT F0.5 weighs precision twice as heavily as recall.
| Proofreader | Precision | Recall | ERRANT F0.5 |
|---|---|---|---|
| GrammarBot Neural | 40.8% | 62.5% | 43.9 |
| AmberPen | 37.4% | 65.5% | 40.9 |
| LanguageTool | 37.8% | 18.3% | 31.2 |
| Sapling | 23.6% | 49.1% | 26.3 |
| Harper 2.4.0 | 5.8% | 13.8% | 6.6 |
GrammarBot Neural leads CWEB F0.5 by 2.99 points. AmberPen has the highest recall, finding 3.6× more of the annotated errors than LanguageTool at similar precision, and remains the strongest product on JFLEG fluency.
Speed
End-to-end latency of each corrector call during the complete benchmark runs, measured per sentence:
| Proofreader | p50 | p95 | p99 |
|---|---|---|---|
| Harper 2.4.0 (local) | 2 ms | 8 ms | 14 ms |
| AmberPen | 217 ms | 1,024 ms | 1,413 ms |
| LanguageTool | 340 ms | 504 ms | 602 ms |
| Sapling | 620 ms | 1,086 ms | 1,646 ms |
| GrammarBot Neural | 1,000 ms | 2,408 ms | 4,421 ms |
Local Harper holds the lowest raw latency. Its in-process timing is not directly comparable to a hosted API call; among hosted correctors, LanguageTool is the most regular. AmberPen is roughly 3× faster than Sapling and 4.6× faster than GrammarBot Neural at the median. Latencies are from the CWEB run (single sentences, ~100 characters); longer inputs take proportionally longer.
Quality, price, and speed
The combined quality score is the arithmetic mean of JFLEG GLEU × 100 and CWEB ERRANT F0.5, giving the two benchmarks equal weight on a 0–100 scale. Higher is better. The shaded upper-left area marks the most attractive quadrant: higher quality at lower cost or latency.
Quality vs price
Entry paid-plan rates. Harper is free and local. LanguageTool is omitted because it is priced per request rather than per character.
Quality vs speed
Median end-to-end latency from the CWEB run. Harper ran in-process; all other products used hosted APIs.
Method
- All correctors received identical inputs; Harper ran locally and the others used their APIs.
- JFLEG scored with the official GLEU scorer and all four references.
- CWEB scored with ERRANT 3.0.2 (spaCy 3.8); gold edits regenerated from CWEB's raw human references with the same tokenizer as the hypotheses.
- AmberPen ran in
correctmode at temperature 0. - Latency percentiles are computed over individual corrector calls, excluding benchmark-runner queue time. Complete runs, July 2026.
Comparisons
- AmberPen vs Harper — quality, local speed, privacy, and language support.
- AmberPen vs LanguageTool — with examples of errors rule-based checking can't see.
- AmberPen vs Sapling — quality, speed, and price.
- AmberPen vs GrammarBot Neural — fluency, CWEB restraint, and latency.