Skip to content
Tutorly

How the AI works

Model-agnostic, and honest about it

Tutorly does not exist without modern language models, and it is not tied to any one lab. Here is the routing, the escalation rule, the fallback chain, and the checks that keep a guiding turn a guiding turn.

Why a model at all

Because teaching is reasoning

A rules engine can check whether an answer is right. It cannot work out why a particular student thinks you can distribute a square across a sum, and it cannot decide which single question would make them notice the error themselves. That judgement lives in understanding both the subject and the person, which is what a language model does.

The harder half is restraint. Any model can print the answer. Holding it back, offering the next small step instead, and confirming only once the student has reasoned there, is the part that teaches, and it is the part Tutorly is built around.

So the model is the tutor. Everything else in this product is plumbing around it: routing to the right model for each move, rendering math cleanly, and keeping a free session honest.

The four steps

What the model is actually asked

01routed to gpt-4.1-mini

Read the attempt

The model takes in the whole exchange: the problem, the student's working, and every wrong turn. A system prompt tells it, firmly, that its job is to guide and never to hand over the answer to graded work.

02routed to gpt-4.1-mini

Find the misunderstanding

It reasons about the specific idea behind the mistake, the wrong belief a good question can correct, rather than just marking the answer wrong. That diagnosis is what decides the next move.

03routed to gpt-4.1-mini

Make one move

A single guiding question, one small step, or an understanding check, sized to where the student actually is. If they are lost, the step gets smaller. If they are close, a nudge. Never a full solution.

04routed to gpt-4.1-mini

Recap and check

At the end, a separate step summarizes what the student worked out, in their own words, and poses one check question so it sticks. Drafting a recap and guiding a turn are different jobs, so they are routed separately.

Routing

The table, generated from the code

This is not a diagram somebody drew. It is rendered from the same routing table the tutoring engine reads, so if it is wrong here it is wrong in production.

StepModelProviderTierUSD per M tokens
Read the attempt and ask the question that moves the student forwardgpt-4.1-miniopenaibalanced$1.60
Write the recap and the understanding checkgpt-4.1-miniopenaibalanced$1.60
Work out the subject, topic and a name for the sessiongpt-4.1-nanoopenaifast$0.40

Escalation

When an exchange gets long, or a misunderstanding will not move after several turns, the guiding step escalates to the frontier model rather than the balanced one, because that is where the harder diagnosis starts to matter more than cost.

Fallback

If the routed model fails or returns nothing, the call walks a chain of candidates from other tiers and other labs before giving up. One provider having a bad afternoon should not end a student's session, and it does not.

Structured turns

Every tutoring turn is validated against a schema before it reaches the student: the message, the kind of move, and whether the answer has been reasoned to. A malformed turn falls back to a gentle prompt rather than showing raw output.

Candidates

What is wired, and what is one key away

Every model below sits behind the same interface. Adding a lab is one case in one file, which is the entire point of building it this way.

gpt-4.1

frontier

openai · 1,000,000 token context

  • multi-step proofs and long derivations
  • a misunderstanding that has not moved after several turns
  • careful writing feedback on a full draft

gpt-4.1-mini

balanced

openai · 1,000,000 token context

  • the guiding turn: a question or a small step
  • reading the student's attempt for the specific error
  • checking understanding before moving on

gpt-4.1-nano

fast

openai · 1,000,000 token context

  • working out the subject and topic
  • naming a session

claude-sonnet

frontier

anthropic · 200,000 token context

  • patient step-by-step guidance
  • encouraging writing feedback

gemini-flash

fast

google · 1,000,000 token context

  • fast turns at scale
  • quick subject triage

llama-open

open

meta · 128,000 token context

  • self-hosted tutoring for schools that cannot send work out

In this deployment the OpenAI and Anthropic adapters are written and the OpenAI one is live. The rest are declared with their real model names and switch on when their key is present. No model is trained on student work, by us or by our providers.