Learn from reality.
Compiler diagnostics, types and working fixes become training examples.
DATAA new model class, in the making
We're building small, local AI models that edit code and check their work against the compiler and tests.
KERN is our first family of VELMs: validated expert language models.
01 / TypeScript prototype Architecture preview
01 / THE MODEL CLASS
A VELM combines a specialized language model with the tools that can check its answers.
For code, those tools already exist: the compiler, the language server and your tests. We're making them part of how KERN learns, improves and responds.
Read the whitepaper (PDF, opens in a new tab)Compiler diagnostics, types and working fixes become training examples.
DATAThe planned next stage rewards code that passes real checks and behavioral tests.
TRAININGThe runtime we're building checks an edit, feeds back failures and verifies the revision.
INFERENCE02 / THE PRODUCT IN MOTION
Fix a type error. Change a function. Repair a failing test. KERN starts with the everyday edits that have a clear way to check the result.
Use the user's name as the display label.
THE VALIDATION LOOP
The model proposes a focused change. The verifier gets the candidate next.
Knowing when to stop is part of the model. We're developing confidence signals so KERN can answer, ask for context or hand off a task that needs a larger model.
03 / LOCAL BY DESIGN
Built for developers and teams who want useful code intelligence on the hardware they already own.
Our current prototype uses a 1.5B-parameter core. The architecture pairs a shared core with small, swappable language experts.
The product is designed to keep code on your machine and work offline, without a cloud inference bill for each edit.
KERN is planned to ship first inside Moxxy (opens in a new tab), our existing open-source CLI and desktop agent harness.
LOCAL RUNTIME + LANGUAGE MODULES ARE IN DEVELOPMENT.
04 / BUILDING IN THE OPEN
parameters in the first
TypeScript prototype
tasks, with hidden tests
and a repeatable harness
pass@1 on 101 answer tasks
mean of two filtered SFT runs
M1 pilot, October 2026. The untuned 1.5B baseline scored 0% on the same answer tasks. These are early, format-specific results, not a production benchmark or frontier-parity claim. See the method and limitations ↗
NOW / PROTOTYPE
Evaluation harness, held-out tasks and the first fine-tuned model. Better data and broader evaluation come next.
NEXT / PRODUCT
Local runtime, verifier feedback, confidence signals and a Moxxy pilot with design partners.
THEN / EXPERTS
Swappable language modules, tested against held-out tasks and the larger-model baselines.
LATER / TEAMS
Larger models for team hardware, on-prem deployments and an API. Release gates come before expansion.
Research targets, not achieved performance. We test each before making a product claim.
70%+of a frontier model’s success rate on scoped code edits, with the same checking budget.
< 3 secresponse time on a laptop. We’ll publish timing by device and edit size.
Know when.Reliable signals to answer, ask for context or hand off beyond the model’s scope.
05 / A FAMILY OF FUTURE EXPERTS
We see VELMs wherever domain knowledge can meet a real check. New domains need the right model core, specialized data and validators built for the work.
EXAMPLE WORKFLOWS
Update the smallest affected function using compiler diagnostics.
Implement a cart total or boundary fix, then run the relevant tests.
WHAT CAN BE CHECKED
Compiler diagnostics, lint, type information and behavioral tests.
TypeScript prototype today. The full runtime loop and Go/Rust experts are in development.
EXAMPLE WORKFLOWS
Draft a reconciliation and flag entries that leave the books unbalanced.
Assemble a filing from source records and check required fields and totals.
WHAT CAN BE CHECKED
Double-entry balance, independently recalculated totals and filing schemas.
Rules need current sources. Accounting treatment and final filings require professional review.
EXAMPLE WORKFLOWS
Map a note into required fields and flag missing or inconsistent information.
Link a draft summary to retrieved clinical sources and surface unsupported statements.
WHAT CAN BE CHECKED
Record schemas, source links, units and explicit consistency rules.
Checks cover specific properties. Clinical interpretation and decisions remain with qualified clinicians.
EXAMPLE WORKFLOWS
Draft a calculation sheet and independently recompute its arithmetic and units.
Flag rule violations or clashes in a model before an engineer reviews it.
WHAT CAN BE CHECKED
Recomputed calculations, unit consistency and structured-model rules.
A validator checks parts of a design. Overall safety and sign-off require a qualified engineer.
EXAMPLE WORKFLOWS
Check that a cited provision exists and matches the relevant date and jurisdiction.
Flag missing required clauses and calculate deadlines from explicit rules.
WHAT CAN BE CHECKED
Authoritative source retrieval, date arithmetic and required-clause checklists.
Legal interpretation is only partially checkable and still needs professional review.
EXAMPLE WORKFLOWS
Draft a SQL change, then check its schema and run it against known fixtures.
Build an analysis script and compare its output with a specified reference.
WHAT CAN BE CHECKED
Executable queries, data schemas, dimensional checks and reproducible outputs.
A reproducible calculation does not establish that the scientific assumptions or conclusions are sound.
EXAMPLE WORKFLOWS
Propose a schedule and flag resource conflicts or capacity overruns.
Match receipts, transfers and shipments against inventory records.
WHAT CAN BE CHECKED
Capacity constraints, inventory conservation and business-rule checks.
Plans need accurate inputs and operational approval. These are future research directions.
Our ambition extends beyond code. These domains are research directions outside the current code product roadmap, with their own core models and validation limits. Select a domain to explore the idea.
LET'S BUILD WHAT COMES NEXT
We're looking for TypeScript teams to help define the first KERN product: the edits that matter, the checks you trust and the way a local model should work.
Become a design partner Early conversations. Direct access to the builders.