Stackness
All moves

Move Attempt: Use a local LLM to fix typos and grammar instead of a cloud model

by algoryunov

LLMs

Most of us now use LLMs to automate our work. Many also want to use local models more. I'm slowly automating my own work too, so I tried a simple "Grammarly-lite" setup on my MacBook (M3 Pro, 18GB). The model fixes grammar and typos and changes the tone of a message, and of course it is important that text never leaves my laptop.

I tried multiple local models, was mainly focused on Qwen3 4B, 8B and 14B. 14B fits in memory, but only just, and of course it is slow, so I mainly measured proofreading with 4B and 8B.

How I tested it:

  • 748 sentences written by English learners. Each one was also corrected by 4 people. I checked how close the model's fixes are to the human fixes.
  • 150 sentences with no mistakes. A good proofreader should not change them.
  • 40 messages to rewrite from casual to formal, or from formal to casual, without losing any facts.
Qwen3-4BQwen3-8B
Grammar score (higher is better)0.6750.718
Fixed exactly like a human did40%44%
Correct sentences left alone65%63%
Tone rewrites with all facts kept92%100%
Time per sentence0.7 s1.2 s

Two examples:

  • "One of this important element is internet." 4B did not change it. 8B fixed it: "One of these important elements is the internet."
  • "The meeting will take place in the Berlin office." I asked for a casual version. 4B wrote "The meeting is in Berlin". The office is gone, so the meaning changed.

My takeaway: use a local model to suggest fixes, not to fix text on its own. It helps clean up a rough draft. But you should not trust it blindly. Both models changed about one in three sentences that had no mistakes. If you accept every change, the model rewrites text that was already fine. Use 8B when quality matters. Use 4B when you want quick suggestions.

Full results, 17 side-by-side examples and how I scored them: https://algoryunov.github.io/llm-local-inference-sheet/site/proofreading.html

Note: I used Qwen3 8B to proofread/fix typos in this move description.

Limits of this test: I used one laptop, and other apps were using a lot of memory, so real speeds may be better. The test sentences are written by learners, so your own text will have fewer and different mistakes. I used one prompt only. A stricter prompt may change fewer correct sentences.

Tools