FixFork
Give FixFork a repo with failing tests. It races three model-written fix attempts on forked branches that can roll back. It returns the branch that passes — with the evidence to prove it.
Runs on Nebius Token Factory with NVIDIA Nemotron models. MIT licensed, source on GitHub.
Watch the demo
A narrated walkthrough of a real run — from the red baseline to the verified winning patch. Under two minutes. Watch on YouTube.
One real run, start to finish
Nothing on this page is staged. A script extracted the numbers below from the run's own artifacts. The raw files are linked at the bottom. This is the FAILED (2 of 2) → green path as the tool recorded it.
- A small Python repo (
examples/demo-repo) ships with a planted bug that makes two tests fail. - FixFork runs the test suite first to confirm the failure — the baseline is red: FAILED (2 of 2).
- Nemotron 3 Super 120B reads the source and the failure output. It proposes three divergent fix hypotheses, each with a concrete edit plan. These are not three variations of one guess.
- The sandbox is forked into three branches; each branch applies its own edit and runs the suite. Branches that still fail get extra repair rounds from the small, fast model (Nemotron 3 Nano 30B).
- The winner is picked by evidence: tests green first, then the fewest changed lines. Losing branches are rolled back — nothing is merged on a hunch.
- The result is exported as a git-applyable patch, a single-file HTML report, and the raw model reply for auditing.
The race real, per-branch tracking
| Branch | Status | Tests | Lines changed | Tokens | Cost |
|---|---|---|---|---|---|
| 1 | green | OK (2 tests) | 1 | 0 | $0.00000 |
| 2 | green | OK (2 tests) | 3 | 0 | $0.00000 |
| 3 | error | 2 | 2850 | $0.00053 |
Tokens and cost are tracked per branch by the router; reasoning-model calls are billed as they run on Token Factory.
Hypotheses raced
- 1. Discount sign inverted — The discount is incorrectly increasing the subtotal because it multiplies by (1 + discount) instead of (1 - discount), causing discounted orders to cost more than the original. (files: src/tax.py)
- 2. VAT applied after discount instead of before — The intended calculation is to apply VAT first, then the discount. The current implementation applies discount (incorrectly) then VAT, leading to an inflated total. Swapping the order fixes the logic. (files: src/tax.py)
- 3. Unnecessary vat_rate parameter shadows global constant — The function defines a vat_rate parameter with default VAT_RATE, but the global constant is sufficient. Keeping the parameter can cause confusion if a different value is passed. Removing the parameter and using the constant directly simplifies the code and eliminates potential misuse. (files: src/tax.py)
The winning fix
diff --git a/src/tax.py b/src/tax.py
--- a/src/tax.py
+++ b/src/tax.py
@@ -6,5 +6,5 @@
def order_total(prices, discount=0.0, vat_rate=VAT_RATE):
"""Total for an order: prices summed, discount applied, then VAT added."""
subtotal = sum(prices)
- discounted = subtotal * (1 + discount)
+ discounted = subtotal * (1 - discount)
return round(discounted + discounted * vat_rate, 2)
The exported patch was re-applied to a pristine checkout with git apply and the test
suite went green (Ran 2 tests ... OK) — download the raw patch and try it yourself.
Web research (Tavily) a later run
This section comes from a later run of the same demo repo, recorded after the run above. FixFork turns the failing-test signature into a Tavily query. It injects the top results into the diagnosis prompt as a hint, not a verdict.
- Query sent:
test_discount_reduces_total AssertionError: 121.0 not less than 100.0 python3 -m unittest discover -s tests - 5 sources returned; top hits: https://github.com/cgoldberg/python-unittest-tutorial · https://github.com/cgoldberg/python-unittest-tutorial/blob/master/README.md and 3 more.
- Same outcome: winner branch 1, 1 line changed, tests green; cost ≈ $0.0037 in model usage (5,071 tokens, all in the diagnosis call).
Full artifacts of this later run: examples/live-run-web-grounded/ on GitHub.
Sandbox verification reproduced twice
The same code ran on the Token Factory Sandboxes backend: opt-in, with VM-level isolation. The full loop ran end to end against a live upstream issue, OpenCTI-Platform/connectors #7778, on 2026-10-01, twice. Each branch forks a content-addressed state image on Nebius infrastructure.
| Run | Baseline (in the VM) | Winner | Patch re-applied | Sandbox ops | Sandbox cost |
|---|---|---|---|---|---|
| A | FAILED (3 of 11) | OK (11 tests) — 7 lines changed | 11 passed (2.18s) | 34/34 | ≈ $0.0259 |
| B | FAILED (3 of 11) | OK (11 tests) — 7 lines changed | 11 passed (1.61s) | 34/34 | ≈ $0.0253 |
Duplicate hypotheses were collapsed automatically: 3 proposed, 2 duplicates dropped, 1 unique edit set kept, with no extra model call. Model usage for these runs: $0.0223 and $0.0235. Case, acceptance criteria and raw logs ship in the repository.
Models & cost
Diagnosis — Nemotron 3 Super 120B
nvidia/nemotron-3-super-120b-a12b
Proposes the three divergent hypotheses from the source and the failing output.
Repair loop — Nemotron 3 Nano 30B
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B
Small, fast loop for extra rounds on branches that still fail — keeps the race cheap.
This run cost ≈ $0.0042 in model usage (7,543 tokens, diagnosis 4,693).
The router detects reasoning-budget exhaustion (finish_reason=length with empty content) and
retries with a doubled budget (4096 → 8192 → 16384 → 32768).
- diagnosis model proposed 3 hypotheses (4693 tokens, $0.0037)
Run it yourself — no API key needed
git clone https://github.com/tuyentran4992/fixfork
cd fixfork
python3 -m fixfork run --repo examples/demo-repo --test "python3 -m unittest discover -s tests -v" --fake --out fixfork-report.md --html
That is the offline demo (deterministic fake router, local sandbox). For a live run,
drop --fake and set NEBIUS_API_KEY to a Token Factory key. Requires Python 3.10+
and git on PATH — FixFork is standard-library only.
Raw artifacts
The files below are exactly what the tool produced for this run:
- Run report (Markdown) — the branch race report
- Run report (HTML) — same report as a single-file page
- Winning patch —
git apply-able - Raw model reply (diagnosis) — three hypotheses as returned