

See what every change did
Prompts and models change weekly. Without numbers, every change is a guess. Every config change is tracked against runs, rework and cost.
Define your agents’ workflow as code, run it on any model and measure every change.

Agents write code in hours. Specifying, verifying and operating it is now the bottleneck.

1. Plan
Agents ship a clear ticket in hours. Specification is now the bottleneck.
Your team’s agent scopes the work and proposes tasks. Nothing starts until you accept.

2. Spec
Agents rarely fail on syntax. They fail on wrong assumptions.
The Spec station writes the spec first and asks when the decision is yours.


3. Build
One general-purpose agent is hard to steer and spends frontier tokens on routine work.
Each station runs its own agent, prompt, tools and model.

4. Review
Review is the new bottleneck: more pull requests, larger diffs.
Spotlight groups each change into reviewable layers, each with the agent’s reasoning.

5. Operate
Shipping faster means more alerts for whoever is on call.
Alerts become tasks. The agent investigates and asks before it touches production.

6. Learn
Generic agents don’t know your dashboards, alerts or conventions, so the same incident happens twice.
Spotlight knows how your team works and the tools it uses. After a fix, it proposes the missing alert, Grafana panel or memory.
Most agent failures come from the harness, not the model. Improving it requires measurement.
7. Improve


Prompts and models change weekly. Without numbers, every change is a guess. Every config change is tracked against runs, rework and cost.

Testing on live work makes your team pay for regressions. Replay a change on past tasks. Keep it only if it wins.
Labs and vendors want to own how you build software. Spotlight keeps it yours.

Agents need your code, credentials and systems. Run them on your servers, in sandboxes like Daytona or E2B, or in our cloud. Use a built-in runner or add your own through an integration.

Every team relies on internal APIs no vendor supports. Connect your tools, or add your own in a few lines.

The best model changes every few months. Choose the harness and model per station.
resource "spotlight_integration" "billing" {team_id = spotlight_team.platform.iddefinition = {integration = "billing-api"actions = {refund = {request = {method = "POST"url = "https://billing.acme.internal/refunds"body = { invoice = "$${{ inputs.invoice }}" }}}}tools = {refund = {does = "refund"description = "Refund an invoice."inputs = { invoice = { type = "string" } }}}}}
An unmanaged harness drifts. Everything is Terraform, reviewed in pull requests.

Agents don’t know how your team does things. Write it down once as a skill, and give it to any station.
Review pull requests and fix failing checks
Pick up issues and report back on them
Answer an agent’s question in the thread
Investigate alerts as they fire
Open work from incidents
Turn errors into tasks
Update status during incidents
Add your own API in a few lines
Open source. Self-host it or use our cloud.