Read the attempt
The model takes in the whole exchange: the problem, the student's working, and every wrong turn. A system prompt tells it, firmly, that its job is to guide and never to hand over the answer to graded work.
How the AI works
Why a model at all
A rules engine can check whether an answer is right. It cannot work out why a particular student thinks you can distribute a square across a sum, and it cannot decide which single question would make them notice the error themselves. That judgement lives in understanding both the subject and the person, which is what a language model does.
The harder half is restraint. Any model can print the answer. Holding it back, offering the next small step instead, and confirming only once the student has reasoned there, is the part that teaches, and it is the part Tutorly is built around.
So the model is the tutor. Everything else in this product is plumbing around it: routing to the right model for each move, rendering math cleanly, and keeping a free session honest.
The four steps
The model takes in the whole exchange: the problem, the student's working, and every wrong turn. A system prompt tells it, firmly, that its job is to guide and never to hand over the answer to graded work.
It reasons about the specific idea behind the mistake, the wrong belief a good question can correct, rather than just marking the answer wrong. That diagnosis is what decides the next move.
A single guiding question, one small step, or an understanding check, sized to where the student actually is. If they are lost, the step gets smaller. If they are close, a nudge. Never a full solution.
At the end, a separate step summarizes what the student worked out, in their own words, and poses one check question so it sticks. Drafting a recap and guiding a turn are different jobs, so they are routed separately.
Routing
This is not a diagram somebody drew. It is rendered from the same routing table the tutoring engine reads, so if it is wrong here it is wrong in production.
| Step | Model | Provider | Tier | USD per M tokens |
|---|---|---|---|---|
| Read the attempt and ask the question that moves the student forward | gpt-4.1-mini | openai | balanced | $1.60 |
| Write the recap and the understanding check | gpt-4.1-mini | openai | balanced | $1.60 |
| Work out the subject, topic and a name for the session | gpt-4.1-nano | openai | fast | $0.40 |
When an exchange gets long, or a misunderstanding will not move after several turns, the guiding step escalates to the frontier model rather than the balanced one, because that is where the harder diagnosis starts to matter more than cost.
If the routed model fails or returns nothing, the call walks a chain of candidates from other tiers and other labs before giving up. One provider having a bad afternoon should not end a student's session, and it does not.
Every tutoring turn is validated against a schema before it reaches the student: the message, the kind of move, and whether the answer has been reasoned to. A malformed turn falls back to a gentle prompt rather than showing raw output.
Candidates
Every model below sits behind the same interface. Adding a lab is one case in one file, which is the entire point of building it this way.
openai · 1,000,000 token context
openai · 1,000,000 token context
openai · 1,000,000 token context
anthropic · 200,000 token context
google · 1,000,000 token context
meta · 128,000 token context
In this deployment the OpenAI and Anthropic adapters are written and the OpenAI one is live. The rest are declared with their real model names and switch on when their key is present. No model is trained on student work, by us or by our providers.