Early access · research tooling for futures strategies

Backtest futures strategies. Believe the results.

VeriRun Lab is a research IDE for futures strategies. Describe an idea in plain English, review the Strategy Brief and the generated code, then run research-grade backtests with conservative fills and an A–F scorecard that isn’t afraid to say NO-GO.

Research tooling, not trading advice.

VeriRun Lab results viewer: a graded backtest with the conservative-fills method banner, per-symbol stat grid, a flagged audit check and an F grade — the reserved holdout not yet evaluated, so no GO exists yet

How it works

From idea to evidence, in four steps

  1. Describe your idea

    Write the strategy the way you’d tell a colleague — “fade the opening range breakout on ES, RTH only.” The builder asks clarifying questions instead of guessing.

  2. Review the Brief & code

    Every assumption is written down in a Strategy Brief you approve first. Generated code arrives as a reviewable diff — production code, tests and test setup accepted separately.

  3. Run it honestly

    Explore fast with an idealized preview, then grade with conservative fills: trade-through required, post-only makers, per-contract fees.

  4. Compare and iterate

    Every run lands in your library with its exact code version, parameters and data slice — graded, comparable and fully traceable.

The AI Strategy Builder: a plain-English idea followed by clarifying questions and a reviewable Strategy Brief The workspace IDE: file tree, Python editor with diagnostics, and run controls The gates tab: pre-committed gates with stored fail verdicts, the benchmark comparison and the fill-realism note — in-sample only, so the grade stays capped until the holdout is evaluated The run library: pinned, titled runs of a VWAP-reversion study with search and filters

The platform

A full research loop, in one place

AI Strategy Builder

Start from a plain-English idea. The builder asks clarifying questions, writes a Strategy Brief with every assumption made explicit, and only then generates code — a multi-file proposal you review as a diff.

The AI Strategy Builder: a plain-English idea followed by clarifying questions and a reviewable Strategy Brief

Clarifying questions before code — nothing applies without your accept.

A real IDE, not a black box

The code is yours to read and change — Python completions, signatures and live diagnostics, with versioned snapshots you can roll back to at any time.

The workspace IDE: file tree, Python editor with diagnostics, and run controls

Every run pins the exact code it ran.

Backtests that tell you the truth

Most backtests flatter you. VeriRun Lab runs an idealized preview to explore quickly, then a conservative pass — trade-through required to fill, post-only maker orders, per-contract fees — before anything gets a grade.

The gates tab: pre-committed gates with stored fail verdicts, the benchmark comparison and the fill-realism note — in-sample only, so the grade stays capped until the holdout is evaluated

A NO-GO verdict caps the grade — the scorecard can’t be talked out of it.

A library of every run

Runs aren’t throwaway. Each one is saved with its grade, settings and results — searchable, taggable and comparable side by side, so last month’s experiment is still evidence today.

The run library: pinned, titled runs of a VWAP-reversion study with search and filters

Share a read-only results link when you want a second opinion.

Provenance for every number

A result you can’t reproduce is a rumor. Every run records the exact inputs that produced it, so any number on screen traces back to its source.

  • Dataset slice & symbols
  • Code version
  • Parameters
  • Engine build

Recorded automatically on every run.

See it before you grade it

The preview chart draws entries and exits on real session structure, so you can eyeball behaviour before spending a conservative run on it.

The preview viewer: a candlestick session chart with entries and exits marked

Idealized preview and conservative grade stay clearly separated.

Study wider, automate further

Optimization studies, one-approval variation fan-out and a multi-instrument screener widen the search; external agents drive the same honest loop over MCP, and the Claude Code team plugin packages the whole workflow.

Same graded scorecards, whichever door you come in through.

Benchmarked against doing nothing

Equity charts overlay buy & hold — on new backtests at your run's own capital with engine costs applied — or a flat-cash floor, so a winning curve has to beat the boring alternatives in plain sight. Older runs show the raw frictionless series and say so.

Excess net and correlation are display stats — they never touch the verdict.

Reality tracking

Link a journal tag to a strategy’s reference backtest and the platform keeps score as live trades come in: In line, Drifting, or Diverged — dated verdicts that append, never a backtest quietly rewritten. The full trading journal sits underneath — imports, MAE/MFE excursions and a real-candle chart for every trade.

Outperforming live never counts against you — only real divergence does.

A leaderboard ranked by evidence

Publishing a run takes audit coverage — a deep audit, robustness checks, a stated hypothesis — and the board ranks by that evidence. Dollar P&L never leaves your account. Alongside it sits a global feedback board where every signed-in user can file bugs and feature requests, vote, and follow status — with a moderation gate keeping it usable.

Ranked by audit evidence, not by return. A modest, robust result outranks a spectacular, fragile one.

Built for research integrity

Guardrails you can feel good about

  • Stored verdicts Reports render the verdict that was saved with the run — never a friendlier recompute.
  • NO-GO included Failing grades are kept and shown like any other result, not quietly discarded.
  • Conservative fills Grades come from the conservative pass: trade-through required, post-only makers, per-contract fees.
  • Train-only scoring Optimization scores the train window only — test is displayed, never scored.

You approve every change

AI proposes; you decide. Generated code lands as a scoped, reviewable diff — gathered in a drafts inbox for review — and nothing is applied to your workspace without an explicit accept.

Strategy code runs sandboxed

Your strategies execute in an isolated sandbox with no network access — runs are reproducible and your machine and data stay out of reach.

Spend stays capped

AI usage runs against your plan's daily quotas — enforced, with honest refusals when a limit is reached. Market-data purchases are quote-first and human-approved — you type the cost to confirm it. No surprise bills, and a clear view of what each session consumed.

Honesty is the default

The grading rubric is versioned and shown with every verdict. Reports carry their method and their caveats — the scorecard can’t be talked out of a NO-GO.

Put your next idea through an honest test

VeriRun Lab is in early access. Open the app and take a strategy from sentence to scorecard.