RESEARCH NOTE / 08 OCTOBER 2026
The evidence
behind the prototype.
M1 is an early TypeScript pilot. Our next research tracks explore 7B and
10B local experts. The pilot gives us a working evaluation harness, a
first fine-tuned 1.5B model and a clear set of questions for the next
stage.
What we measured.
The filtered v0.3 runs individually scored 35.6% and 32.7% pass@1.
Their mean, 34.2%, is the number on the landing page. The unfiltered
runs scored 38.6% and 32.7%.
The tuned 1.5B model outperformed the prompted 7B baseline in this
protocol. It remains well below the 32B baseline. This is not evidence
of frontier parity or a general coding leaderboard result.
The method and its limits.
-
Golden TS v0.1 contains 120 held-out tasks: 101 code-answer tasks
and 19 clarify/refuse tasks. Training data is kept separate from the
golden set.
-
Code answers are evaluated by the versioned TypeScript harness
against compilation, lint, format and hidden behavioral tests. A
passing compile alone is not a successful answer.
-
The first SFT experiments use a Qwen2.5-Coder-1.5B core and
synthetic examples from GLM-5.2. The v0.3 datasets contain 5,565
training records per variant.
-
The pilot is small. Seeds differ by roughly 3 to 6 percentage
points, and baselines were each run once. We need more tasks,
repeated runs and a human audit before drawing broader conclusions.
-
The filtered versus unfiltered comparison did not confirm the
expected quality benefit of toolchain filtering. The paired
bootstrap interval for the pass@1 difference was -7.4 to +4.5
percentage points. Format compliance improved, but answer accuracy
did not.
-
The test suite bounds what is validated. Hidden tests reduce the
chance of a superficial fix, but cannot establish correctness for
every possible input.
THE NEXT MODEL SIZES / RESEARCH PLAN
From 1.5B to 7B and 10B.
The 1.5B pilot established the workflow and exposed a capacity gap on
harder edits. Our next research direction is to test larger cores
while preserving the same domain specialization, local deployment goal
and evaluation discipline. The final base family is still being
selected.
7B
NEXT SCALING EXPERIMENT
More capacity for harder edits.
Evaluate a roughly 7B core as the next candidate for a local
KERN expert. The aim is better function edits, refactors and use
of diagnostics than the 1.5B prototype. The wider base-model
comparison includes the 4B to 9B class; 7B is a research track,
not a frozen architecture.
- Training plan
-
Improve and deduplicate the data, compare LoRA with full
fine-tuning, then test distillation and toolchain-reward
training. Introduce each stage separately so its effect can
be measured.
- Local target
-
Quantized inference on 16 GB-class laptops, with the model,
language module, KV cache and developer tools measured
together. A memory fit alone does not establish usable
latency.
- Release gate
-
Beat the 1.5B prototype and the untuned 7B baseline on
held-out tasks, with repeated runs and acceptable accuracy
after quantization. The baseline result in the table is not
a trained KERN-7B result.
A stronger expert, still close to the work.
Explore a roughly 10B model for scoped, single-language edits.
The whitepaper proposes either a dense core or a
mixture-of-experts model with about 3B active parameters. We
want to test how close a specialized local model can get to much
larger models when both have the same verifier loop.
- Architecture options
-
Compare dense and sparse designs, shared language modules
and 4-bit inference. About 6 GB of quantized weights is a
planning estimate, before KV cache, modules and runtime
overhead.
- Training plan
-
Continued pretraining on code and diagnostics, supervised
edits, on-policy distillation and reinforcement learning
with compiler and behavioral-test rewards. Confidence probes
and faster decoding are separate experiments.
- Release gate
-
Evaluate TypeScript, Go and Rust separately, including
results with and without the feedback loop. Use the same
attempt budget and checks for KERN and larger baselines, and
publish paired uncertainty intervals and device timing.
7B and 10B are planned research tracks, not released checkpoints. We
have not demonstrated their quality, memory use or speed.
Whole-application generation and repository-wide autonomous tasks are
outside this scoped-edit target. Base selection and native multi-token
prediction support remain open decisions.
First local measurements.
Five preliminary samples on a Mac M4 with 24 GB RAM, using MLX and the
filtered v0.3 1.5B model in bf16, reached 24 to 28 tokens per second
and about 3.4 GB peak memory. Warm time to first token was 0.6 to 0.7
seconds; the cold first call took 1.6 seconds.
Complete 70 to 90-token answers took 3.5 to 4 seconds. Answers of 230
to 260 tokens took about 11 seconds. The sub-3-second goal has not
been reached. Quantized inference, broader devices and p50/p95 latency
remain to be measured. Five samples on one device are a preliminary
check, not a latency guarantee.
What still has to be built.
A product-ready local runtime and broader device-level latency
evaluation; the inference-time compile/fix loop; reinforcement
learning with toolchain rewards; reliable confidence probes; and the
shared-core, swappable-language product architecture. Go and Rust are
planned experts, not released modules.
The original research goals include at least 70% of a frontier model's
editing success rate and sub-3-second responses on a laptop. These are
targets to test, not achieved performance. A larger task set, fair
tool budgets and device-specific timing are necessary to evaluate
them.
The VELM idea beyond code.
A specialized model coupled to a domain validator is the broader
research direction. Accounting can use balanced ledgers and
recalculated totals. Engineering can use explicit rule checks. Medical
workflows could use structured-record validation, reference retrieval
and consistency checks, while clinical interpretation still needs
expert review.
These are exploration areas outside the current code product roadmap.
A new domain needs a suitable model core, domain-specific data and a
validator with a clearly defined scope. It is not simply a released
adapter for the current code model. There is no available KERN
medical, accounting or engineering model today.
Architecture and research plan.
The English whitepaper describes the VELM principle, proposed
architecture, evaluation gates and longer-term research directions.
Its performance and size projections are hypotheses. The M1 results
above are the current measured pilot evidence.
Read the whitepaper ↗ (PDF, opens in a new tab)