Skip to content

Changelog

RightNow Agent releases

Earlier releases sit below the checkpoint.

RightNow Agent 1.0.5

2026-09-09

Windows console pastes stay in one chip, and the model list refreshes after connecting. Release fixes keep Windows ARM64 downloads consistent with the published binary.

What changed

FIXED

  • Multi-line pastes in the Windows console (ConPTY) fold into one paste chip.
  • The /model list refreshes when the catalog arrives after connecting.
  • Windows ARM64 artifacts use the publish-signed executable with a refreshed checksum and runtime evidence.

IMPROVED

  • RIGHTNOW_INPUT_TRACE=1 prints raw input events for paste diagnostics.
  • The release runbook includes the first-alpha checklist.
  • The release runbook includes the first-stable checklist with observed signed evidence.

RightNow Agent 1.0.4

2026-09-09

The first hosted release brings native installers and in-app updates to Windows, macOS and Linux. Six platforms. Linux binaries publish GPG signatures, and every binary lists its SHA-256 in the release manifest.

What changed

NEW

  • Native binaries for Windows x64, Windows ARM64, macOS Apple Silicon, macOS Intel, Linux x64 and Linux ARM64.
  • Per-platform binaries have SHA-256 sidecars and a schema-2 manifest, with cli/stable and cli/alpha channel pointers.
  • install.ps1 for Windows and install.sh for macOS and Linux verify downloads against the sidecar and manifest.
  • Native installers use ~/.rightnow/bin by default, keep .old rollbacks, install completions, update PATH and record installer.toml.
  • The in-app updater uses the hosted channel and verifies the manifest, sidecar and size before activation.
  • The TUI honours --disallowed-tools as well as headless runs.

IMPROVED

  • The shipped version 6 configuration uses the RunInfra OpenAI-compatible endpoint, RIGHTNOW_API_KEY and qwen3-8-flash-next.
  • Local runtime paths for Ollama, llama.cpp, local weights, GPU inference and local embeddings are removed.
  • Rust config sync is additive and fail-closed: it retains owner values and refuses invalid, newer or concurrently changed files.
  • One shell.turn.speed_shadow diagnostic event records model calls, tool rounds and estimated savings per turn.

RightNow Agent 1.0.3

2026-09-05

The legacy Windows x64 zip release installs under %LOCALAPPDATA%\Programs\RightNow. It is superseded by the native installer introduced in 1.0.4.

What changed

NEW

  • Legacy Windows x64 zip distribution.
  • Installation under %LOCALAPPDATA%\Programs\RightNow.
  • Superseded by the native installer introduced in 1.0.4, with legacy files left intact.

CHECKPOINT

2026-09-05

The editor line ends here. From this date on, every entry is RightNow Agent, the coding agent.

Editor 1.0.0Feb 4, 2026
Agents, Skills, and More Languages, Editor 1.0.0

Agents, Skills, and More Languages

Create custom agents, extend capabilities with skills and MCPs, and develop GPU kernels in CUDA, Triton, Mojo, PyTorch, Numba, and more.

What changed

NEW

  • Custom agents with skills and MCP integrations for GPU workflows
  • Clear model selection across cloud LLMs and local GPU-backed models
  • Numba support with native docs, autocomplete, emulation, profiling, and benchmarking
  • CUDA Tile support with native docs, autocomplete, emulation, profiling, and benchmarking
  • Mojo support with native docs, autocomplete, emulation, profiling, and benchmarking

IMPROVED

  • More stable SSH and remote GPU workflows after the migration
  • Updated chat and agent sessions experience on the new core
Forge 0.1.0Jan 5, 2026
Introducing Forge, Forge 0.1.0

Introducing Forge

CLI Swarm Agent for generating production-ready CUDA/Triton kernels. Up to 5x faster than torch.compile() with 97.6% correctness rate.

What changed

NEW

  • 32 Parallel Coder+Judge Agent Pairs: Swarm architecture generates and validates kernels concurrently
  • MAP-Elites Evolutionary Optimizer: 36 behavior cells with 4 specialized islands (memory_bound, compute_bound, fused_ops, tensor_cores)
  • Pattern RAG System: 1,711 CUTLASS patterns + 113 Triton patterns for context-aware generation
  • Up to 5x Speedup: Llama-3.1-8B (5.2x), Qwen2.5-7B (4.2x), Mistral-7B (3.4x), SDXL UNet (2.9x) vs torch.compile
  • Multiple Input Types: HuggingFace model IDs, KernelBench tasks (250+), or custom PyTorch files
  • Dual Output Formats: Triton (Python GPU kernels) or native CUDA C++ - drop-in PyTorch replacement
  • Three Optimization Modes: --turbo (fast ~2min), default (balanced), --quality (maximum optimization)
  • Interactive CLI: forge command launches wizard, plus forge browse for KernelBench task browser
  • Session Management: Track past optimizations with forge session list
  • Credit System: 1 credit per KernelBench/custom kernel, 1-2 credits for HuggingFace models

IMPROVED

  • 97.6% Correctness Rate: Tiered evaluation pipeline (Dedup → Compile → Test → Benchmark)
  • 250k Tokens/Second: Powered by fine-tuned NVIDIA Nemotron 3 Nano 30B for fast generation
  • Automatic Tensor Core Optimization: WMMA, TMA for Hopper architecture
  • Cross-platform Install: npm, npx, curl (macOS/Linux), PowerShell (Windows)
Editor 0.2.0Jan 1, 2026
PyTorch Kernel Support, Editor 0.2.0

PyTorch Kernel Support

Profile, benchmark, and emulate PyTorch kernels directly in the editor. Same workflow as CUDA, Triton, TileLang, and CUTE.

What changed

NEW

  • PyTorch Kernel Profiling: Profile custom PyTorch kernels with full NCU integration
  • PyTorch Benchmarking: Run statistical timing analysis on PyTorch operations
  • PyTorch Emulation: Test PyTorch kernels across 86+ GPU architectures without hardware
  • Automatic Kernel Detection: Detects PyTorch kernels from your code automatically
  • Cross-DSL Comparison: Compare PyTorch kernel performance side-by-side with CUDA/Triton implementations
  • Unified Workflow: Same profiling methods (NCU Full, Fast, Static, Line-by-Line) work across all supported languages
Editor 0.1.0Dec 8, 2025
Multi-DSL GPU Development Platform, Editor 0.1.0

Multi-DSL GPU Development Platform

RightNow AI now supports Triton, TileLang, and CUTE alongside native CUDA, with intelligent documentation retrieval that understands your GPU and code context.

What changed

NEW

  • Multi-DSL Platform: CUDA, Triton, TileLang & CUTE support with automatic detection and compilation
  • Unified Profiling Experience: Profile Triton and TileLang kernels with NCU just like CUDA kernels
  • CUTLASS Auto-Detection: Automatically finds your CUTLASS installation across common locations
  • Enhanced Language Features: Full semantic highlighting, hover documentation, and go-to-definition for all DSLs
  • Context-Aware Help: AI automatically retrieves relevant GPU documentation based on your code and questions
  • GPU-Specific Filtering: Only shows documentation compatible with your GPU architecture
  • 100+ Documentation Sources: Comprehensive coverage of CUDA, Triton, TileLang, and CUTE APIs
  • DSL-Aware Suggestions: AI understands which DSL you're working with and provides relevant guidance
  • GPU Context Display: See your active GPU info directly in the chat
  • DSL-Specific Metrics: Extracts Triton warps/stages, TileLang block sizes, CUTE tile dimensions
  • Smart Performance Warnings: Get alerts for high warp counts, excessive sync points, or register pressure

IMPROVED

  • Benchmark Accuracy: Triton benchmarks now show correct DSL columns (BLOCK, Warps, Stages)
  • Emulation Reliability: Fixed bug where Triton used real GPU instead of emulator when emulated GPU selected
  • GPU Detection: Better architecture identification and compatibility checking
  • Chat Interface: Cleaner message rendering with improved code display

FIXED

  • Triton Emulation: Now correctly uses emulation mode when emulated GPU is selected
  • Benchmark Columns: Triton and TileLang show proper DSL-specific columns instead of CUDA defaults
  • Parameter Handling: DSL parameters now correctly flow through profiling pipeline
Editor 0.0.76Nov 30, 2025
Mac Support & Multi-GPU Profiling, Editor 0.0.76

Mac Support & Multi-GPU Profiling

Full macOS compatibility with Metal GPU detection for Apple Silicon. Multi-GPU profiling to compare GPU vs GPU side-by-side.

What changed

NEW

  • Mac Platform Support: Full macOS compatibility with Metal GPU detection for Apple Silicon (M1/M2/M3) and Intel GPUs
  • Multi-GPU Support: Select and profile multiple GPUs simultaneously with side-by-side performance comparison
  • GPU Filter Dropdown: Filter profiling results by specific GPU
  • Remote GPU Support: Full SSH connection to remote GPU servers with seamless CUDA environment integration
  • Dynamic Profiling Methods: 5 profiling options (NCU Full, Fast, Static, Line-by-Line, Kernel Replay) with quick-switch buttons

IMPROVED

  • Simplified Chat Modes: Streamlined to Agent, Gather, and Forge only (Coming soon)
  • Chart Rendering: Better performance visualization for profiling data

FIXED

  • Double Credit Consumption: Fixed bug charging credits twice
  • Profiling Data Persistence: No more data loss on view refresh
  • SSL Handshake Issues: Fixed connection stability

BREAKING

  • Iterate Mode: Removed and consolidated into simplified chat mode system
Editor 0.0.45Oct 30, 2025
Execution-Driven Emulator & Agentic AI Optimization, Editor 0.0.45

Execution-Driven Emulator & Agentic AI Optimization

Cycle-accurate GPU emulation with 96-98% accuracy. No physical GPU required. AI automatically iterates and optimizes kernels to peak performance.

What changed

NEW

  • New GPU Emulator built from scratch with cycle-accurate scheduling and multi-warp latency simulation
  • PTX and SASS translation for deep low-level analysis and debugging
  • Remote connection support (SSH + WSL) to run and profile kernels anywhere
  • Agentic AI "Iterate Mode" that writes, profiles, and optimizes kernels automatically until peak performance
  • Kernel fusion detection for better performance across sequential and parallel operations (Beta Users)

IMPROVED

  • Enhanced benchmarking and profiling with full metric breakdowns and bottleneck detection
  • Local LLMs now work perfectly
  • 96-98% emulator accuracy vs real GPUs
  • Simulation speed under 100ms for 1,000 instructions
  • 30% more accuracy than previous builds
Editor 0.0.31Sep 22, 2025
Remote GPU Access & AI Insights, Editor 0.0.31

Remote GPU Access & AI Insights

Code anywhere, run everywhere. Connect to remote GPUs with SSH and cloud providers.

What changed

NEW

  • Remote GPU connection via SSH integration
  • Native support for GPU cloud providers (RunPod, Google Cloud, AWS, Azure, Paperspace, Vast.ai, Lambda Labs)
  • Seamless profiling on remote GPUs as if they were local
  • Automatic GPU detection on remote machines
  • Smart Profiling Terminal with AI-powered insights
  • Automatic bottleneck detection (memory-bound vs compute-bound)
  • NCU-compatible metrics without requiring hardware
  • AI-generated optimization suggestions (memory coalescing, bank conflicts, occupancy, branch divergence)

IMPROVED

  • Fixed NCU GUI integration for report generation
  • Enhanced profiling UI with collapsible sections
  • WebWorker-based analysis for non-blocking performance
  • LRU caching for instant re-analysis
  • Improved error handling and fallback mechanisms
Editor 0.0.30Sep 18, 2025
Full GPU Emulator - No Hardware Required, Editor 0.0.30

Full GPU Emulator - No Hardware Required

Profile any CUDA kernel without a physical GPU. Choose from 86+ GPU architectures.

What changed

NEW

  • Full GPU emulator for profiling without physical hardware
  • 86+ GPU architectures supported
  • Static kernel analysis engine (under 100ms)
  • Roofline model implementation with ±15% accuracy
  • Architecture comparison across multiple GPUs instantly
Editor 0.0.29Sep 14, 2025
Benchmarking Terminal & Static Profiling, Editor 0.0.29

Benchmarking Terminal & Static Profiling

Full benchmarking terminal with visual kernel comparisons and instant CodeLens insights.

What changed

NEW

  • Benchmarking Terminal for benchmark sweeps and custom kernel configurations
  • Visual comparison between kernels
  • Static Profiling with instant CUDA kernel insights in CodeLens
  • Real-time registers, shared memory, and occupancy analysis while typing
  • Profile with Configs - complete cycle with persistent configs and history
  • Tools Detector (nvidia-smi, nsight compute, nvcc)
Editor 0.0.28Sep 14, 2025
CUDA Benchmarking System, Editor 0.0.28

CUDA Benchmarking System

Comprehensive benchmarking with execution time, memory bandwidth, occupancy, and multi-GPU support.

What changed

NEW

  • Execution time, memory bandwidth, occupancy, SM efficiency, and register usage metrics
  • Data size presets, warmup runs, and execution controls
  • Grid/block optimization with automatic suggestions
  • Multi-GPU support with device-specific benchmarking
  • Session management with persistence across restarts
  • Sortable results with performance indicators
  • CSV export for sharing benchmark results
Editor 0.0.20Aug 18, 2025
Multi-LLM Provider Support, Editor 0.0.20

Multi-LLM Provider Support

Support for 15+ AI providers including local models.

What changed

NEW

  • OpenAI, Anthropic, Deepseek integration
  • Local Ollama and vLLM support
  • Provider configuration flexibility
  • Fill-in-the-Middle autocomplete
Editor 0.0.10Aug 5, 2025
Initial Release, Editor 0.0.10

Initial Release

First public release of RightNow AI.

What changed

NEW

  • NVIDIA Nsight Compute integration
  • Real-time GPU performance metrics
  • Hardware detection and optimization
  • CUDA syntax highlighting and IntelliSense