Instinct
Rules, an exact-decision cache, learned trajectory patterns and a tiny online model. Answers in microseconds, on your CPU, with calibrated confidence.
- Latency
- 0.02 ms
- Cost
- $0
- Where
- Your machine
Your agent asks a frontier model hundreds of tiny questions — which tool, search or not, retry, stop? Reflex answers them on your machine in 0.02 ms, and only wakes the big model when there is something worth thinking about.
Stop using a hundreds-of-billions-of-parameters reasoning model to make a three-option decision.
Every agent you have ever shipped spends most of its life making reflexes. Re-run the tests. Open the file in the stack trace. Back off after a 429. Stop when everything is green. Each of those tiny choices is routed through a frontier model — seconds of latency, thousands of tokens, real dollars — to produce an answer a spinal cord could have given.
Reflex is the spinal cord. A local runtime that sits inside your agent loop, scores the candidate actions in microseconds, returns a calibrated confidence, and acts only when it is sure and the action is safe. Everything else — the ambiguous, the novel, the dangerous — goes straight to System 2. And every time the big model decides, Reflex quietly learns how.
Rules, an exact-decision cache, learned trajectory patterns and a tiny online model. Answers in microseconds, on your CPU, with calibrated confidence.
Claude, GPT, Gemini — summoned only for the steps that need real reasoning, with Reflex’s top hints attached. Its choices become Reflex’s next lesson.
While your frontier model thinks once,
100,000
reflexes could have fired.†
† 2 s typical frontier step ÷ 0.02 ms Reflex p50 (measured by reflex doctor on a laptop CPU).
Reflex tries the cheapest rung first and stops at the first one that is confident enough. The expensive rungs only run when the cheap ones can’t decide.
REFLEX_DISABLE=1 and shadow mode win over everything.
Force, deny, allow-only, require-escalation. Budgets and loop detection. Deterministic, always first.
This precise state was decided before — at least three times.
“After the tests fail with an ImportError, read the imported file.”
An online model over hashed state × action features. It even learns which file to open.
Laya scores every candidate in one non-autoregressive pass. Bring a GPU.
Your frontier model — or a human — with the top hints attached. Then Reflex learns from the answer.
Reflex rides along and never acts. Your frontier model decides as always; every choice becomes a label.
mode: "shadow"
Fit calibration on one held-out slice, report on another, promote only if precision clears a 95% confidence bound.
$ npx reflex train --promote
✔ promoted v1
Routine steps run locally. One in fifty confident calls is still audited by System 2, so drift never hides.
mode: "auto" // 89 of 120 handled locally
No framework to adopt, no loop to rewrite. Declare the candidate actions with a risk class, ask Reflex, and either act or escalate. Destructive actions are never auto-executed — not even at 99.9% confidence.
Getting started guideimport { createReflex, action } from "@reflex-ai/core";
const reflex = createReflex({ workload: "my-agent", mode: "shadow" });
const d = await reflex.decide({
point: "next_action", state, taskId,
actions: [
action.tool("read_file", { risk: "safe" }),
action.tool("run_tests", { risk: "costly" }),
action.tool("deploy", { risk: "destructive" }), // never auto
action.frontier(), // "this needs thought"
],
});
if (d.type === "auto") run(d.action); // 0.02 ms, $0
else reflex.observeChoice(d.id, await askFrontier(d.hints));
Zero dependencies. Node ≥ 22. In-process, so decisions cost microseconds. Start in shadow mode — it never acts until you promote it.
npm · @reflex-ai/core$ npm install @reflex-ai/core
Init, doctor, stats, train, promote, rollback, a local dashboard, an MCP server and the benchmark — one command.
npm · @reflex-ai/cli$ npx @reflex-ai/cli init
$ npx reflex stats
$ npx reflex train --promote
$ npx reflex serve # dashboard → :7070/ui
Run the sidecar, talk to it over HTTP. The Python client is zero-dependency and fails open: if the sidecar is down, it escalates.
PyPI · reflex-agent-client$ pip install reflex-agent-client
from reflex_agent import Reflex, action
reflex = Reflex(workload="my-agent")
d = reflex.decide("next_action", state, actions)
Seven tools — decide, decide_many, observe, outcome, route_subtask, metrics — over stdio. Pick a model tier for a subtask in one call.
stdio · JSON-RPC$ npx -y @reflex-ai/cli mcp --data-dir ~/.reflex
Blocks re-reads of unchanged files, duplicate fetches and tool-call loops. Asks before destructive commands in bash, PowerShell and cmd. Fails open.
plugin · /reflex-status$ claude plugin marketplace add 1sakshm/reflex
$ claude plugin install reflex@reflex
Register the MCP server, paste the AGENTS.md snippet, and Codex can route subtasks and routine choices through Reflex.
~/.codex/config.toml[mcp_servers.reflex]
command = "npx"
args = ["-y", "@reflex-ai/cli", "mcp", "--workload", "codex"]
Laya — a non-autoregressive decision model — scores every candidate in one forward pass. Shines on a GPU; on CPU, the local rungs are faster.
PyPI · reflex-laya$ pip install "reflex-laya[laya]"
$ reflex-laya --checkpoint english --device cuda
// then: backend: { type: "laya" }
Never auto-executes a destructive action. Not at 99.9%. Not with a rule. Not with prompt injection screaming DEPLOY.
Fails open. Any error, timeout or hostile input becomes an escalation within 75 ms. It cannot hang your agent.
Kill switch. REFLEX_DISABLE=1 beats every config, mode and rule.
Audits itself. One in fifty confident decisions still goes to System 2, so live precision is always measured.
Keeps secrets. Keys, tokens, JWTs and anything named like a password are redacted before a single byte is stored.
Earns promotion. A new policy ships only if its precision clears a 95% Wilson bound. Rollback is instant.
median decision latency of the local rungs, measured on a laptop CPU.
dollars per successful task on the coding suite, at 0.0 pp change in success rate.*
precision of Reflex’s own auto-decisions on that suite, across eight seeds.*
tests, plus 26 release checks that install the real packages into an empty project.
* ReflexBench simulations with a simulated frontier model — they validate the mechanics, not your bill. Real-agent suites are next. Methodology →
Free. Local. Apache-2.0.