A coding agent for your terminal, built for one thing: speed. It ships while other harnesses are still warming up.
Where the milliseconds go
A staged turn at demo pace, with the measured durations. Hover a step to read what it cost.

Session open in 105 to 115 ms, not 631 to 728 ms cold

Bounded edits can finish from the tool result, with no narration call

1,976 tokens instead of 9,676. Model call 35% faster

SHA-256 receipts and exit codes your scripts can trust
One turn, from prompt to receipt

Bring the model provider you already use
RunInfra ships in the profile. The others are OpenAI-compatible entries you add with your own base URL, model slug, and key.
Cerebras Inference. A small edit task measured as a one-second turn.
GroqCloud LPU inference.
Our own hosted Model APIs at api.runinfra.ai. The default profile.
Any OpenRouter model, one key.
Vercel AI Gateway routing through one endpoint.
Where the harness gets out of the way

Session open in 105 to 115 ms
The client stays warm between sessions. Harness overhead measured 269 to 286 ms with the daemon reused, against roughly 1,100 ms with no daemon.

1,976 prompt tokens, not 9,676
Only the tools the turn needs are sent. 80% fewer tokens and a 35% faster model call.

No narration call when the postcondition verifies
A bounded edit can complete from the tool result and the file on disk, with no narration-only model call.

Narration is not evidence
Verified changes get SHA-256 receipts. Exit codes 0, 20, 21, 22 and 23 say exactly what happened.
Latest updates and improvements
New stable SSH, custom agents with skills and MCP, clear cloud/local LLM list, plus Numba, CUDA Tile, and Mojo support with docs, autocomplete, emulation, profiling, and benchmarking.
CLI Swarm Agent with 32 parallel Coder+Judge pairs. Up to 5x faster than torch.compile() with 97.6% correctness.
Profile, benchmark, and emulate PyTorch kernels directly in the editor. Same workflow as CUDA, Triton, TileLang, and CUTE.
CUDA, Triton, TileLang & CUTE support with intelligent documentation retrieval that understands your GPU and code context.

Install RightNow Agent
One command in your terminal, then sign in with your RightNow account.
Windows x64
For Windows 11/10 x64
Windows ARM
For Windows 11/10 ARM
Mac Apple Silicon
For M1, M2, M3, M4 Macs
Mac Intel
For Intel-based Macs
Linux x64
Debian, Ubuntu, Fedora
Linux ARM
ARM64 distributions