Which local model works best with Proselon?
A model running on your own computer is the one writing engine that costs nothing, asks for no account, and sends no sentence of your book anywhere. It is also the hardest of the engines to get a good experience from, and the reason is narrow enough to state at the outset. Proselon's co-writer works by calling tools, and most models small enough to run on an ordinary machine cannot call tools well enough to finish a piece of work.
What the co-writer requires of a model
Proselon's co-writer is not a chat window with your manuscript pasted into it. It opens the files your chapters are stored in, reads them, searches across the book for a name or a thread, and writes its changes back to disk. Each of those actions is a tool call, and the model issues them itself, many times over in the course of a single request. A model that cannot call tools in the OpenAI style cannot operate the co-writer, however well it writes a paragraph on its own. Where Proselon can recognise that failure it says so in one line — the model is not compatible with the co-writer, and another one is the cure — rather than handing you the error the server produced.
The second requirement is room. Before it drafts a scene the co-writer reads a great deal: your Template's workflow, the book's canon, the samples of your own prose, the plan for the scene, and the scenes that come before it. Proselon requires a context window of at least 128K tokens to hold all of that, and the app states the figure in the settings panel. Hosted engines carry more than that as a matter of course. Local servers do not. Ollama starts a model in a window of a few thousand tokens unless the context length is raised in its settings, and whatever exceeds the window is dropped with no error anywhere, so the co-writer writes a scene having never seen your book and the prose thins for no visible reason. Proselon reads the window where the server reports it, tells you the number it found, and says which setting to raise when the number is short.
What we have tested
The list is short. We would rather publish a table of models we have run than a longer one assembled from what models claim about themselves.
| Model | Size | Works with Proselon | Notes |
|---|---|---|---|
| Qwen 3.5 9B | 9B | Yes | Completes co-writer work in our testing. The smallest model we can presently point a writer toward. |
| Qwen 3 8B | 8B | Yes | Completed a full co-writer turn through a custom server address, July 2026. Tested as a working connection rather than across a whole book. |
| Mistral 7B | 7B | No | Its own chat template rejects the co-writer's request before the model reads any of it. Measured August 2026. No setting corrects this. |
Last updated September 2026. The table will grow as we test more models, and a model's absence from it means we have not run it rather than that it fails.
Setting one up
- Install Ollama or LM Studio and fetch a model that reports tool support.
- Start the server. In LM Studio that is the Developer tab; Ollama serves whatever you have installed in it as soon as the app is open.
- Raise the context length to 128K or beyond. In Ollama it is a setting on the server. In LM Studio you choose it as the model loads.
- In Proselon, open Settings, then Writing Engine, and choose Local Model.
- Under Running with, pick Ollama, LM Studio, or Other for any other server that speaks the OpenAI format. Each of the first two fills in its standard Server address for you, which you can change if yours differs.
- Watch the row under the address. It reports what Proselon finds there, whether the server is running and how many models it holds, and it refreshes every few seconds, so starting the server in another window shows up here without a restart.
- Choose your model under Model. Models the server is holding in memory appear under Loaded and the rest under Not loaded, greyed and unselectable.
That last division is worth a word, because an installed model is not a model that can answer. Ollama pulls an unloaded model into memory on first use, which is minutes of silence, and LM Studio refuses outright. Loading the model in the program that serves it moves the row up into the list you can choose from.
Choosing a model we have not tested
Both programs mark which of their models advertise tool support, and either mark is the place to begin. Ollama's model library carries a Tools filter, and every model listed under it wears a tools label of its own, so the filtered library is a shortlist in itself. LM Studio puts a hammer badge on a model with native tool use, as its documentation on tool use describes, and notes that a model without native support falls back to a default with variable results. Our Mistral 7B measurement is what that variability amounts to in practice, which is why a badge is a shortlist rather than a guarantee.
Size is the other consideration, and it is a question about your computer rather than about the model. A model runs at a workable speed when it fits in the memory available to it, which on an Apple Silicon Mac is the unified memory the whole machine shares and on a Windows machine is the memory on the graphics card. One that does not fit will still run by falling back on ordinary system memory, at a pace that makes drafting a chore. A larger model is generally more capable and slower, so the one to reach for is the largest your machine holds comfortably.
Whether it is worth the trouble
For most writers, a signed-in engine gives a better experience for less work, and setting one up takes a minute. The local route earns its difficulty when something particular is true of you: you write where there is no connection, or you have exhausted the limits of every plan you are willing to pay for, or your manuscript is of a kind you would rather no company's servers ever held. A capable computer is a precondition of it rather than a preference.
We have not yet found a model small enough to run on most computers that also does what the co-writer needs of it. That is a report on this year rather than a permanent state of affairs, and the table above is where any change to it will appear first.
Proselon