KERNBack to the product

RESEARCH NOTE / 08 OCTOBER 2026

The evidence
behind the prototype.

M1 is an early TypeScript pilot. Our next research tracks explore 7B and 10B local experts. The pilot gives us a working evaluation harness, a first fine-tuned 1.5B model and a clear set of questions for the next stage.

What we measured.

Model / run Answer pass@1 Compile rate
Qwen2.5-Coder 1.5B, untuned 0.0% 15.8%
Qwen2.5-Coder 7B, untuned 10.9% 39.6%
Qwen2.5-Coder 32B, untuned 66.3% 80.2%
KERN 1.5B, v0.3 filtered SFT (2-seed mean) 34.2% 67.8%
KERN 1.5B, v0.3 unfiltered SFT (2-seed mean) 35.6% 68.8%
Golden TS v0.1, 101 answer tasks. Data taken from the repository's M1 report, dated 8 October 2026. Baselines use the KERN response format through prompting; SFT models are trained on it.

The filtered v0.3 runs individually scored 35.6% and 32.7% pass@1. Their mean, 34.2%, is the number on the landing page. The unfiltered runs scored 38.6% and 32.7%.

The tuned 1.5B model outperformed the prompted 7B baseline in this protocol. It remains well below the 32B baseline. This is not evidence of frontier parity or a general coding leaderboard result.

The method and its limits.

THE NEXT MODEL SIZES / RESEARCH PLAN

From 1.5B to 7B and 10B.

The 1.5B pilot established the workflow and exposed a capacity gap on harder edits. Our next research direction is to test larger cores while preserving the same domain specialization, local deployment goal and evaluation discipline. The final base family is still being selected.

7B

NEXT SCALING EXPERIMENT

More capacity for harder edits.

Evaluate a roughly 7B core as the next candidate for a local KERN expert. The aim is better function edits, refactors and use of diagnostics than the 1.5B prototype. The wider base-model comparison includes the 4B to 9B class; 7B is a research track, not a frozen architecture.

Training plan
Improve and deduplicate the data, compare LoRA with full fine-tuning, then test distillation and toolchain-reward training. Introduce each stage separately so its effect can be measured.
Local target
Quantized inference on 16 GB-class laptops, with the model, language module, KV cache and developer tools measured together. A memory fit alone does not establish usable latency.
Release gate
Beat the 1.5B prototype and the untuned 7B baseline on held-out tasks, with repeated runs and acceptable accuracy after quantization. The baseline result in the table is not a trained KERN-7B result.
10B

LARGER LOCAL EXPERT

A stronger expert, still close to the work.

Explore a roughly 10B model for scoped, single-language edits. The whitepaper proposes either a dense core or a mixture-of-experts model with about 3B active parameters. We want to test how close a specialized local model can get to much larger models when both have the same verifier loop.

Architecture options
Compare dense and sparse designs, shared language modules and 4-bit inference. About 6 GB of quantized weights is a planning estimate, before KV cache, modules and runtime overhead.
Training plan
Continued pretraining on code and diagnostics, supervised edits, on-policy distillation and reinforcement learning with compiler and behavioral-test rewards. Confidence probes and faster decoding are separate experiments.
Release gate
Evaluate TypeScript, Go and Rust separately, including results with and without the feedback loop. Use the same attempt budget and checks for KERN and larger baselines, and publish paired uncertainty intervals and device timing.
7B and 10B are planned research tracks, not released checkpoints. We have not demonstrated their quality, memory use or speed. Whole-application generation and repository-wide autonomous tasks are outside this scoped-edit target. Base selection and native multi-token prediction support remain open decisions.

First local measurements.

Five preliminary samples on a Mac M4 with 24 GB RAM, using MLX and the filtered v0.3 1.5B model in bf16, reached 24 to 28 tokens per second and about 3.4 GB peak memory. Warm time to first token was 0.6 to 0.7 seconds; the cold first call took 1.6 seconds.

Complete 70 to 90-token answers took 3.5 to 4 seconds. Answers of 230 to 260 tokens took about 11 seconds. The sub-3-second goal has not been reached. Quantized inference, broader devices and p50/p95 latency remain to be measured. Five samples on one device are a preliminary check, not a latency guarantee.

What still has to be built.

A product-ready local runtime and broader device-level latency evaluation; the inference-time compile/fix loop; reinforcement learning with toolchain rewards; reliable confidence probes; and the shared-core, swappable-language product architecture. Go and Rust are planned experts, not released modules.

The original research goals include at least 70% of a frontier model's editing success rate and sub-3-second responses on a laptop. These are targets to test, not achieved performance. A larger task set, fair tool budgets and device-specific timing are necessary to evaluate them.

The VELM idea beyond code.

A specialized model coupled to a domain validator is the broader research direction. Accounting can use balanced ledgers and recalculated totals. Engineering can use explicit rule checks. Medical workflows could use structured-record validation, reference retrieval and consistency checks, while clinical interpretation still needs expert review.

These are exploration areas outside the current code product roadmap. A new domain needs a suitable model core, domain-specific data and a validator with a clearly defined scope. It is not simply a released adapter for the current code model. There is no available KERN medical, accounting or engineering model today.

Architecture and research plan.

The English whitepaper describes the VELM principle, proposed architecture, evaluation gates and longer-term research directions. Its performance and size projections are hypotheses. The M1 results above are the current measured pilot evidence.

Read the whitepaper (PDF, opens in a new tab)