Valta Docs
Enforced proxy (OpenAI)
Calling cap.allow() from your own code is cooperative: code that skips the check can still reach OpenAI. The enforced proxy removes that option. The agent never holds your OpenAI key — it holds a Valta virtual key, which only works against Valta.
agent (virtual key) → Valta /v1/chat/completions → Cap + freeze check
→ denied / can't decide? stop. OpenAI is not called.
→ approved? Valta → OpenAI (your stored key) → response
It uses the same Cap settings, ledger, deny reasons, and audit trail as allow() — enable Cap and set limits on the agent's page exactly as you would for allow().
Setup
New to this? Follow the step-by-step guide — it includes a deny test that needs no OpenAI credit and a troubleshooting table. The short version:
- On the agent's page (
/dashboard/agents/<agentId>), Enable Cap and set limits. - In Enforced proxy (OpenAI) on the same card, Save key (your OpenAI key, stored encrypted; only the last 4 characters are ever shown) and Create virtual key (shown once — copy it).
- Point any OpenAI client at Valta:
OPENAI_BASE_URL=https://valta.co/v1
OPENAI_API_KEY=vk_live_... # Valta virtual key, not an OpenAI key
from openai import OpenAI
client = OpenAI() # reads OPENAI_BASE_URL and OPENAI_API_KEY
client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Summarize this ticket."}],
max_tokens=300, # bounds the pre-call estimate
extra_headers={"X-Valta-Run-Id": "ticket-4812"}, # groups calls for the per-run limit
)
The same setup over the API
Full request/response details: Proxy API reference.
# Store the agent's OpenAI key (encrypted at rest; response carries last4 only)
curl -X POST https://valta.co/api/v1/agents/support-bot/provider-keys \
-H "x-api-key: $VALTA_API_KEY" -H "Content-Type: application/json" \
-d '{"provider":"openai","apiKey":"sk-proj-..."}'
# Mint a virtual key — "secret" is returned once
curl -X POST https://valta.co/api/v1/agents/support-bot/virtual-keys \
-H "x-api-key: $VALTA_API_KEY" -H "Content-Type: application/json" \
-d '{"name":"production"}'
# Revoke one
curl -X DELETE https://valta.co/api/v1/agents/support-bot/virtual-keys/vk_... \
-H "x-api-key: $VALTA_API_KEY"
Provider keys and virtual keys are live-mode only; a test-mode Valta API key gets a 400.
What happens on each call
- Virtual key is checked (unknown or revoked →
401). - Plan and rate limit for that key (see Plans below).
- Cost estimate from the request: prompt size (~4 characters per token) plus
max_tokens/max_completion_tokensas the output bound, at the model's OpenAI list price. Nomax_tokens→ Valta assumes 1,024 output tokens. Models Valta doesn't have a price for are estimated at the most expensive tier, so they're denied early rather than under-counted. - Cap decision — freeze, Cap enabled, per-run / daily / monthly limits, plan limits. Same checks as
allow(). - Approved: Valta sends one request to
https://api.openai.com/v1/chat/completionswith your stored key and returns OpenAI's response unchanged. Valta does not retry. - Settle: the ledger is set to the real cost from OpenAI's
usagenumbers — up or down. If OpenAI rejects the request (4xx), the estimate is released. If OpenAI times out or 5xx's, the estimate stays charged, since OpenAI may have billed.
Approved responses carry X-Valta-Allow-Id, X-Valta-Run-Id, X-Valta-Estimated-Usd, and X-Valta-Actual-Usd.
Denies — no request reaches OpenAI
| Status | reason | Meaning |
|---|---|---|
| 402 | per_run_limit, daily_limit, monthly_limit | The agent's Cap limit would be exceeded. |
| 402 | plan_limit, plan_agent_limit | Your Valta plan's monthly tracked spend or agent count is used up. |
| 402 | plan_key_limit | This virtual key is over your plan's virtual-key allowance (e.g. after a downgrade). |
| 403 | frozen | The agent is frozen. |
| 403 | cap_disabled | Cap isn't enabled for this agent, or it has no policy. |
| 403 | no_provider_key | No OpenAI key is stored for this agent. |
| 429 | rate_limited | Too many requests on this virtual key in the last minute. |
| 503 | valta_unavailable | Valta couldn't make a decision (e.g. its database is unreachable). Fails closed. |
Body — OpenAI-compatible error plus flat fields to match on:
{
"approved": false,
"reason": "per_run_limit",
"id": "allow_…",
"message": "Valta per-run limit reached for this run. No request was sent to OpenAI.",
"error": { "message": "…", "type": "valta_denied", "code": "per_run_limit" }
}
Plan denies (plan_limit, plan_agent_limit, plan_key_limit) also carry plan, limit, upgradeTo, and upgradeUrl, and the message says which limit was hit and where to upgrade. For example, a Free account past its $50:
{
"approved": false,
"reason": "plan_limit",
"id": "allow_…",
"message": "Your Valta Free plan's tracked spend for this month ($50) is used up. No request was sent to OpenAI. Upgrade to Builder ($500 tracked per month) at https://valta.co/dashboard/subscriptions. Tracked spend resets on the 1st of the month (UTC).",
"plan": "FREE",
"limit": "$50 tracked per month",
"upgradeTo": "Builder",
"upgradeUrl": "https://valta.co/dashboard/subscriptions",
"error": { "message": "…", "type": "valta_denied", "code": "plan_limit" }
}
The agent's page shows the most recent refusal under Last deny in the proxy panel.
OpenAI SDKs don't auto-retry 402 or 403, so a deny can't turn into a retry storm.
Run ids
Send X-Valta-Run-Id to put several calls in one run (an agent loop, a job, a conversation) so the per-run limit applies across them. Without it, each call is its own run, and the per-run limit acts as a per-call limit.
The run's spend is kept by Valta, not by your process. If the agent crashes and a new process picks up the same run id with the same virtual key, the run is still charged for everything already spent — the next call over the cap is 402 per_run_limit and never reaches OpenAI.
Plans
Proxied spend counts toward the same Cap plan budget as allow(). Minting a key past the limit returns 402 with code: "PLAN_VIRTUAL_KEY_LIMIT" and the upgrade link.
| Free ($0) | Builder ($29/mo) | Startup ($99/mo) | Enterprise | |
|---|---|---|---|---|
| Active virtual keys (per account) | 1 | 10 | 50 | Unlimited |
| Requests per minute, per key | 10 | 120 | 600 | 3,000 |
| Cap agents | 1 | 10 | 50 | Unlimited |
| Tracked USD per month | $50 | $500 | $5,000 | Unlimited |
You pay OpenAI for tokens, on your own key; Valta never resells or marks them up. The Valta plan is what keeps the check in front of every call. Going over the plan's tracked spend is a 402 plan_limit deny, not a silent pass.
Limits today
- OpenAI Chat Completions only.
stream: truereturns a 400. - Use
https://valta.co/v1as the base URL (notwww.valta.co). - If OpenAI itself rejects an approved call (for example
insufficient_quotaon an unfunded key), its error is passed through and the estimate is released. - Tracked cost uses OpenAI list prices; it can differ slightly from your OpenAI invoice (cached input, batch, or negotiated pricing aren't modeled).
- An approved call without
max_tokenscan cost more than its estimate. The real cost is recorded afterwards and denies the next call — setmax_tokensfor a strict per-call bound.
Health
GET https://valta.co/v1/health returns status: "ok", or "degraded" (HTTP 503) when Valta can't reach its database — in which case every proxied call is being denied, not forwarded.
Example
Valta-hq/langgraph-spend-cap — a LangGraph retry loop that never holds an OpenAI key. Hops 1–3 show up on your OpenAI usage; hop 4 is denied by Valta and never does.