Nebius x NVIDIA Global AI Hackathon · Coding & Agentic Engineering track

FixFork

Give FixFork a repo with failing tests. It races three model-written fix attempts on forked branches that can roll back. It returns the branch that passes — with the evidence to prove it.

Runs on Nebius Token Factory with NVIDIA Nemotron models. MIT licensed, source on GitHub.

Run it yourself GitHub repository

failed / 2baseline: FAILED (2 of 2)
3hypotheses raced as branches
branch 1winner — 1 line fix
≈ $0.0042run cost · 7,543 tokens

Watch the demo

A narrated walkthrough of a real run — from the red baseline to the verified winning patch. Under two minutes. Watch on YouTube.

One real run, start to finish

Nothing on this page is staged. A script extracted the numbers below from the run's own artifacts. The raw files are linked at the bottom. This is the FAILED (2 of 2) → green path as the tool recorded it.

  1. A small Python repo (examples/demo-repo) ships with a planted bug that makes two tests fail.
  2. FixFork runs the test suite first to confirm the failure — the baseline is red: FAILED (2 of 2).
  3. Nemotron 3 Super 120B reads the source and the failure output. It proposes three divergent fix hypotheses, each with a concrete edit plan. These are not three variations of one guess.
  4. The sandbox is forked into three branches; each branch applies its own edit and runs the suite. Branches that still fail get extra repair rounds from the small, fast model (Nemotron 3 Nano 30B).
  5. The winner is picked by evidence: tests green first, then the fewest changed lines. Losing branches are rolled back — nothing is merged on a hunch.
  6. The result is exported as a git-applyable patch, a single-file HTML report, and the raw model reply for auditing.

The race real, per-branch tracking

BranchStatusTestsLines changedTokensCost
1greenOK (2 tests)10$0.00000
2greenOK (2 tests)30$0.00000
3error22850$0.00053
Winner: branch 1 — branch 1 passed the tests (OK (2 tests)) with 1 line(s) changed

Tokens and cost are tracked per branch by the router; reasoning-model calls are billed as they run on Token Factory.

Hypotheses raced

The winning fix

diff --git a/src/tax.py b/src/tax.py
--- a/src/tax.py
+++ b/src/tax.py
@@ -6,5 +6,5 @@
 def order_total(prices, discount=0.0, vat_rate=VAT_RATE):
     """Total for an order: prices summed, discount applied, then VAT added."""
     subtotal = sum(prices)
-    discounted = subtotal * (1 + discount)
+    discounted = subtotal * (1 - discount)
     return round(discounted + discounted * vat_rate, 2)

The exported patch was re-applied to a pristine checkout with git apply and the test suite went green (Ran 2 tests ... OK) — download the raw patch and try it yourself.

Web research (Tavily) a later run

This section comes from a later run of the same demo repo, recorded after the run above. FixFork turns the failing-test signature into a Tavily query. It injects the top results into the diagnosis prompt as a hint, not a verdict.

Full artifacts of this later run: examples/live-run-web-grounded/ on GitHub.

Sandbox verification reproduced twice

The same code ran on the Token Factory Sandboxes backend: opt-in, with VM-level isolation. The full loop ran end to end against a live upstream issue, OpenCTI-Platform/connectors #7778, on 2026-10-01, twice. Each branch forks a content-addressed state image on Nebius infrastructure.

RunBaseline (in the VM)WinnerPatch re-appliedSandbox opsSandbox cost
AFAILED (3 of 11)OK (11 tests) — 7 lines changed11 passed (2.18s)34/34≈ $0.0259
BFAILED (3 of 11)OK (11 tests) — 7 lines changed11 passed (1.61s)34/34≈ $0.0253

Duplicate hypotheses were collapsed automatically: 3 proposed, 2 duplicates dropped, 1 unique edit set kept, with no extra model call. Model usage for these runs: $0.0223 and $0.0235. Case, acceptance criteria and raw logs ship in the repository.

Models & cost

Diagnosis — Nemotron 3 Super 120B
nvidia/nemotron-3-super-120b-a12b
Proposes the three divergent hypotheses from the source and the failing output.

Repair loop — Nemotron 3 Nano 30B
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B
Small, fast loop for extra rounds on branches that still fail — keeps the race cheap.

This run cost ≈ $0.0042 in model usage (7,543 tokens, diagnosis 4,693). The router detects reasoning-budget exhaustion (finish_reason=length with empty content) and retries with a doubled budget (4096 → 8192 → 16384 → 32768).

Run it yourself — no API key needed

git clone https://github.com/tuyentran4992/fixfork
cd fixfork
python3 -m fixfork run --repo examples/demo-repo --test "python3 -m unittest discover -s tests -v" --fake --out fixfork-report.md --html

That is the offline demo (deterministic fake router, local sandbox). For a live run, drop --fake and set NEBIUS_API_KEY to a Token Factory key. Requires Python 3.10+ and git on PATH — FixFork is standard-library only.

Raw artifacts

The files below are exactly what the tool produced for this run: