Introducing Sys1

Your coding agent is a slow, careful thinker. Sys1 gives it a fast one for the small decisions, so it spends fewer tokens and less time on work that doesn’t need deep thought.

Release at publication: v0.17.0. Install from GitHub with Bun. Review and verification are experimental and advisory; local Qwen is experimental, and TypeSafe’s hosted Jev model is opt-in.

Watch a coding agent work and you will see it spend real effort on small questions. Did that test fail because of my change? Does this diff break the project’s rule about error handling? Did I actually push? Each one costs a frontier model context and seconds, and the answer is usually yes or no.

Psychologists call fast, instinctive thinking System One and slow, deliberate thinking System Two. A coding agent is all System Two. Sys1 adds the other half: it sends those small questions to Jev, TypeSafe’s hosted System One decision model, which returns a probability in under a second. TypeSafe’s published workflow figures put a decision at about $0.0004. The agent keeps its context for the work that needs it.

Sys1 ships as project skills for Codex, Claude Code, and Devin, plus a CLI and an API for your own code. This post explains how the pieces fit, where ALGAL comes in, and what a first small trial did and did not show.

A 52-second introduction to the workflow. Original motion graphics and instrumental score, rendered for Sys1 with Slopcamera. Read the visual transcript.

An agent, a fast model, and a language that remembers

Sys1 is one part of a three-part system. The agent is the large language model you already use. It plans, writes code, and works through the problems that need deliberate reasoning.

Jev handles the reflexes. A question with a declared answer shape (yes or no, a choice among named options, or a score) comes back as probabilities your code can act on, not a paragraph to parse. Sys1 routes the question, checks the answer, and hands the agent a result it can use in one step.

ALGAL is a programming language for agent programs. An ALGAL program can wait for a person’s approval, resume after a crash, and replay what it did from its receipts. A procedure that proves itself on declared test cases is kept with its evidence and reused. That is what lets a harness get better with use instead of starting from scratch every session.

Sys1 already works this way on a small scale. When a review turns up a real mistake, you can draft a repository rule from it, and every later review checks for it. Feedback on each finding is kept, and reviews of unchanged code are reused instead of paid for twice. ALGAL takes the same idea further, to whole procedures. Today, Sys1 and ALGAL are separate open-source projects that call the same Jev decision API.

What a decision looks like

In TypeSafe’s System One approach, an application supplies state and questions, and a model returns structured decisions with probabilities. The caller defines what the possible answers mean. A question can ask whether a condition holds, choose among named alternatives, or score evidence against ordered criteria.

Jev is TypeSafe’s hosted model for those decisions. Sys1 is the layer around the call. It offers a Jev-compatible API, picks the backend you configured, checks the answer against the question that was asked, and reports which backend answered.

That distinction matters when choosing where to run a decision. Sys1 can send a request to hosted Jev, an installed Qwen model, or a compatible HTTP service you configure. Enabling Jev selects hosted-only routing. An outage returns an error instead of silently substituting a local model. The local adapter approximates the answer format with generic Qwen models; shared JSON shapes do not establish equal quality or calibration.

The API calls its three answer types noul, choice, and score. These names describe the data your code receives. Validation can reject an unknown choice or a malformed probability, but the caller still needs evidence that the model makes useful judgments on its own tasks. The Jev setup guide explains the connection, and the runtime design traces the request.

Start with the evidence an agent can inspect

The review workflow begins with a Git change and a rule. A repository might require a login path to finish storing a token before reporting success. A review question can bring that requirement together with the selected diff. The answer points the agent toward a candidate to investigate in the implementation and its tests.

The sys1-review skill supplies the sequence: preview a checkpoint, inspect findings, and record whether each was useful, incorrect, or unverifiable. You choose the paths and rules. The bundled rules cover newly empty catch blocks and removed test assertions in JavaScript and TypeScript; broader language or policy coverage needs suitable rules and examples.

A review checkpoint

  1. SelectDiff + rule

    Choose the Git changes and the requirement to check.

  2. PreviewEvidence to send

    Inspect selected files, skipped evidence, and request count.

  3. EvaluateSelected model

    Send the prepared question; validate the returned answer.

  4. InvestigateCandidate finding

    Read the code and tests; record useful, incorrect, or unverifiable.

Before the final message

sys1 verify compares claimed outcomes with reachable Git, pull-request, and page evidence.

Agreement, contradiction, or unavailable evidence
The model contributes a judgment. The agent’s investigation and the repository’s required checks determine the next action.

This workflow also keeps costs down. A preview makes no model calls. An unchanged batch is reused for up to 24 hours instead of being paid for again, and findings the agent has already seen stay out of new output. A checkpoint reports skipped evidence, so a partial review is visible. When a repair changes the evidence, a fresh checkpoint evaluates that change; a changed file alone does not establish that a reported problem was fixed.

The separate sys1-verify skill checks a proposed completion message. Git state can contradict a claim that all work was committed or pushed. A linked pull request can contradict a claimed merge. A fetched page can be compared with a claim that a change is live. Missing evidence stays unverifiable. File or stdin input works with any agent; check-command results are available through local Devin transcript discovery.

Both skills install project instructions without adding automatic hooks or activating a backend. Their usefulness depends on the repository’s rules and the selected model. An earlier review experiment missed its reproduced defect and flagged two of five control hunks, which is why findings require investigation and ordinary tests remain part of the work.

The trial points to one next experiment

Sys1 also includes source-checkout profiles for failure triage, excerpt relevance, and claim support. A profile fixes the questions and model route while accepting new input. On September 28, 2026, Jev 1.13.0 received 48 synthetic examples across these profiles, each submitted once. It returned 45 valid answers matching the authored labels and three errors.

All 48 submitted examples

16 per workflow · Jev 1.13.0 · September 28, 2026

  • Matched label
  • Wrong baseline label
  • Request error

Failure triage

Jev16 / 16
Error signatures16 / 16

Both matched every label. These examples show no advantage over ordinary error matching.

Excerpt relevance

Jev13 / 16 3 errors
Token overlap6 / 16

Every valid answer matched its label, but three submissions failed. Error handling needs work before a broader trial.

Claim support

Jev16 / 16
Literal matching6 / 16

A reason to test more representative claims against a stronger comparator. Literal matching is a deliberately weak baseline.

Each bar groups 16 results by outcome. Eight development examples and eight separately authored screening examples were frozen for each workflow. Both sets are synthetic and public. Errors remain in the denominator. Full report and observations.
Read the chart as a table
Matched labels and errors across all submissions
WorkflowJev matchedJev errorsBaseline matched
Failure triage16/16016/16
Excerpt relevance13/1636/16
Claim support16/1606/16

Claim support is the most useful next candidate from this small screen. Its baseline predicted insufficient evidence on every example and matched six labels, so this comparison leaves stronger entailment models and better rules untested. The profile judges a claim against the supplied excerpt; it does not verify that the excerpt itself is true.

The relevance errors also matter. Two responses reached the gateway with provider HTTP 200 and failed response validation. The first failure was a gateway HTTP 502 whose underlying cause was not retained. The original failed request was never retried; a continuation submitted only the 30 previously unsubmitted cases. These observations test the profiles and evaluator, not the installed review skill or production accuracy.

This trial measured answers, not savings. Per-decision cost and latency come from TypeSafe’s published figures, and the separate compact check output skill returned 35% less text across 563 real test and build runs. Nobody has yet measured how many tokens a whole agent task saves; that is the next number worth getting.

The profile evaluator makes a similar comparison available from a source checkout. It validates examples offline before you opt into calls, then records coverage, matched labels, errors, timing, and routes. A next trial can preserve the first result, use stronger baselines, and measure how the decision changes the completed task.

Try a checkpoint on a change you understand

After installing Sys1, start inside a Git repository with a small change and a rule you can inspect. Install the project instructions and preview a staged checkpoint. The preview lists the planned work without calling a model.

Preview selected source and rules
sys1 review setup codex --dry-run --json
sys1 review setup codex
sys1 review checkpoint --staged --model typesafe/jev-1.13.0 \
  --max-requests 10 --dry-run --json -- src test

Use setup claude-code or setup devin for those agents. To evaluate with Jev, provide TYPESAFE_API_KEY privately in your environment, run sys1 jev enable, and repeat the checkpoint without --dry-run. The selected source goes to that backend. The review guide walks through investigation and feedback; the verification guide adds the final-message check.

If your application already has a decision to make, the Node/Bun client and HTTP API expose the same question types. Begin with labeled examples from that task and compare the result with the code you would otherwise use.