Episodic Reset Contracts for API-Backed RL Environments

Stateful API simulators enable reinforcement learning by enforcing deterministic episode resets.

Staff Writer · · 11 min read
Cover illustration for “Episodic Reset Contracts for API-Backed RL Environments”
RL Environment Design · October 9, 2026 · 11 min read · 2,446 words

Episodic reinforcement learning depends on a clean boundary between rollouts, a boundary that collapses the moment the environment is a live third-party API rather than a simulator you control, and in a conventional simulation, resetting an episode is a function call: the physics engine zeroes out, the grid reinitializes, the game state returns to frame zero. When the environment is Stripe, Slack, or GitHub, reset means something much harder: undoing durable side effects and restoring state deterministically before the next rollout begins. There is no gradient-based undo for a Stripe charge that has already processed or a Slack message that has already landed in a channel. Once the agent takes that action, the consequence is permanent in a way no backward pass can touch.

This creates a problem that goes beyond inconvenience. Without a defined reset contract, the agent's next observation depends on the residue of whatever happened in the prior rollout, and that dependency violates the Markov property that policy gradient algorithms assume. Methods like PPO work only if each episode's state transitions are drawn from a consistent, well-defined distribution, independent of what some other rollout did five minutes earlier. When a training run's "environment" is really just an account on a live API, with real records accumulating across thousands of parallel containers, that independence quietly disappears, and nothing in the training loop will tell you that it did. Training requires specifying the reset problem in advance, with a precise definition of what "reset" means for that API. Hoping the agent's actions won't collide, or cleaning up by hand after the fact, is not a substitute for that specification.

What a reset contract must specify

A reset contract is a machine-checkable specification, and it has to answer four questions before a single rollout can begin. First, it defines the canonical starting state of every resource the agent's action space can touch, not just the resources it is expected to touch on the happy path. Second, it defines the procedure for getting back to that starting state from whatever residue the prior episode left behind. Third, it lists the invariants that must hold true before the first action of each rollout is taken. Fourth, it defines the session scope that keeps one rollout's mutations from bleeding into another's.

The starting state clause means something only when it reflects a realistic environment. An empty Stripe account with no customers, no subscriptions, and no invoices doesn't give an agent anything representative to act on. A design-studio Stripe world, pre-populated with existing customers, live subscriptions, and a history of invoices, gives the agent a starting context that actually resembles production. Stateful API simulators, Rystic among them, ship with exactly this kind of pre-seeded state built in, giving the reset contract a concrete, production-like baseline to specify and check against.

A simplified contract, sketched as a schema, might look like this:

reset_contract:
  starting_state:
    resources: [customers, subscriptions, invoices, webhooks]
    seed: "design_studio_v3"
  procedure:
    idempotent: true
    scope: session
  invariants:
    - no_charges_in_status: [processing, disputed]
    - no_unread_channel_messages
    - webhook_delivery_journal: empty
  session_scope:
    key: session_id
    isolation: atomic

The reset procedure itself has to be idempotent: calling it twice from two different post-episode states has to land on identical starting conditions every time, or the contract isn't really a contract. It also can't require restarting the server process, because parallel rollouts share infrastructure, and a global flush would stall every other container's training run. The invariant set is what turns the contract from a document into something enforceable: no charges sitting in a non-initial status, no channels carrying unread messages from a prior episode, a webhook delivery journal that comes back clean. A health-check endpoint or assertion suite runs these checks before the rollout is permitted to start, so a broken reset gets caught before it contaminates training data, not after. And the session scope clause assigns each rollout a session key that namespaces its mutations, so reset(session_id) is an atomic operation scoped to just that rollout. If you skip this part, parallel rollouts end up sharing one global state namespace, so you get correlated observations instead of the independent, identically distributed rollouts the training algorithm needs.

Why stateless mocks cannot enforce a reset contract

A stateless mock is disqualified from serving as an RL training environment because the category itself cannot represent the state transitions a reset contract has to specify and verify. Mocks are still genuinely useful for unit tests, where the point is to check that a client handles one known response shape correctly. That is a different job from simulating an environment an agent explores over many steps.

A mock server maps request patterns to fixed responses. It has no internal memory, so it cannot express the sequence that actually matters for training an agent: create an order, fetch it back, cancel it, fetch it again, and see a different status reflecting the cancellation. If an agent's policy depends on recognizing a state transition, a payment moving from processing to succeeded to chargeback, a stateless stub will pass a broken implementation silently.

Mocks also can't fire webhook events triggered by state changes, because webhooks are consequences of state transitions and a stateless server has no state to transition. Many of the actions an agent most needs to learn, confirming a payment capture, confirming email delivery, watching a prediction market settle, communicate their outcome through an asynchronous webhook. An environment that can't simulate webhook delivery can't teach an agent to wait for and act on that kind of confirmation.

Stateless mocks also drift faster than stateful simulators because they're hand-maintained: every change to the real API requires someone to go in and manually edit the stub, and the drift typically becomes visible only once it appears in production. If you check simulators continuously against the live API, you catch that drift before release, and the reset contract's assumptions stay in sync with what the real service actually does. Finally, the reset contract requires a scoped reset(session_id) operation, and that concept is undefined for a stateless server, which has no session state in the first place to scope or clear.

What stateful simulators add as the minimum viable substrate

A stateful simulator is the minimum viable substrate for an API-backed RL environment, because only this layer can represent cumulative state, fire event-driven side effects, and expose a scoped reset operation, the three things a reset contract needs to be enforced.

With cumulative state, every call sees the ones that came before it. An added customer appears when the customers are listed. A deleted customer is gone from the list when it's queried again. That sequence satisfies the Markov property at the layer where the agent actually interacts, rather than something a test harness has to reconstruct from configuration after the fact. A session-based state machine tracks each client's position within a multi-step flow, so the same GET /v1/payments/{id} call can correctly return processing, then succeeded, then chargeback, depending on what happened earlier in that same session.

Event-driven side effects follow from the same state machine. Timed webhooks fire between states, signed the way the real API signs them, complete with retries and a delivery journal, so the agent can learn to wait for and respond to asynchronous confirmation in a way a stateless mock has no mechanism to teach. Business-logic faults, a chargeback that arrives asynchronously, a KYC rejection, a webhook signature the agent never bothers to validate, depend on this same state machine rather than on a fault-injection decorator bolted onto a fixed response.

And scoped reset means winding the state machine back to episode-start conditions without restarting the server process, so parallel rollouts can share infrastructure without one container's reset corrupting another's in-progress episode. Rystic's simulators ship with a cohesive pre-seeded world for each API they cover: a mid-sized design studio populated across Slack, PayPal, Lob, Resend, and Stripe, a large open-source repository standing in for GitHub, and recorded Kalshi markets for Kalshi. So that pre-seeding turns the canonical starting state the contract specifies into something concrete you can use right away, rather than something a team has to build from nothing.

Fidelity verification as a first-class contract requirement

Diagram: Simulator Fidelity Scores by API. Visualizes: Show the agreement rates Rystic publishes for each API simulator versus the live API, so readers can immediately see which simulators are tightest and which have the most drift.

A stateful simulator satisfies the structural requirements of a reset contract, but structure alone isn't enough. A simulator that isn't checked continuously against the live API will drift silently, and a policy trained against a drifted simulator ends up optimizing for a reward landscape that no longer matches the real service it's meant to prepare the agent for.

The gap here comes from how APIs actually change over time. Endpoints get added, response shapes shift, and status-code semantics get revised, so a simulator maintained by hand will always lag a step behind those changes. For a reset contract specifically, drift means the canonical starting state the contract defines no longer matches what the real API would actually produce; this invalidates the guarantee the contract exists to provide.

Continuous verification addresses this by running the same set of requests against both the live API and the simulator before every release and measuring how often the two agree. Rystic publishes these agreement rates per API: Slack at 93.9% across 294 probes, GitHub at 98.2% across 327 probes, Kalshi at 91.7% across 361 probes, PayPal at 100.0% across 89 probes, Lob at 100.0% across 24 probes, Okta at 96.9% across 159 probes, Linear at 87.7% across 81 probes, and Resend at 81.0% across 42 probes. A published fidelity score like this is a machine-readable signal that a given simulator version is fit to serve as the environment for a given batch of episodes, and it can be treated as a precondition on the contract, checked in CI before a training run is allowed to start.

Kimi's application-layer simulators for Gmail, Notion, Slack, and Canvas are owned entirely by the team that builds them, which maximizes control over reset behavior but severs the link to how the real service currently behaves, trading fidelity for convenience. For evaluating an agent's behavior in a fixed scenario, that's a reasonable design choice. For training an agent that will eventually act against the real API, the simulator can now diverge from reality indefinitely at a higher level of abstraction, with no external check catching the divergence.

Fault injection as an episodically consistent contract clause

A reset contract that only specifies the happy-path starting state is incomplete, because real agents have to learn to handle failure, and fault injection that fires at random introduces noise into the reward function rather than into the environment's actual transition dynamics. A noisy reward function destabilizes policy gradient learning in ways indistinguishable from ordinary exploration difficulty, which makes random fault injection worse than no fault injection at all for training purposes.

Two layers of fault injection operate differently and shouldn't be collapsed into one concept. Network and protocol faults, latency injection, 429 rate-limit responses, 503 service-unavailable errors, get injected at the transport layer and can be applied deterministically per episode if the contract specifies them precisely: this rollout class sees a 429 on the third call, for instance. Business-logic faults are a different animal. A payment that succeeds and then charges back, a KYC rejection that arrives asynchronously, a webhook signature the agent never validates, each of these depends on what already happened earlier in the episode rather than on a random draw made at call time, so expressing them requires a genuine stateful state machine.

For RL training, the requirement that matters most is episodic consistency: a fault scenario has to be reproducible across rollouts, or the policy gradient has no stable signal to learn from. That means the reset contract has to specify which fault class applies to a given batch of rollouts, and the simulator has to be able to arm that fault scenario as part of the reset operation itself, not fire it off asynchronously at some unpredictable point later. A fault that appears consistently on the fifth call within a session is something a policy can learn to anticipate and route around; a fault that appears on a random call with no connection to session history is just noise dressed up as difficulty. Rystic's Pro tier exposes fault injection, covering 429s, 503s, and latency, alongside settable state. This lets both the starting-state clause and the fault clause of a reset contract be armed together in the same reset call rather than configured through two disconnected systems.

Staggered resets and the parallel-safety requirement at scale

Everything above describes what a single reset contract has to specify. At scale, with thousands of parallel training containers running rollouts simultaneously, the way resets are scheduled across that fleet becomes its own requirement, separate from what any one reset contains.

Synchronous batch resets, where every container resets at the same moment, produce correlated state across rollouts. That correlation skews the learning signal and destabilizes training in ways that look like environment difficulty but are really artifacts of how the reset schedule was structured. Staggered Environment Resets Improve Massively Parallel On-Policy Reinforcement Learning (arXiv:2511.21011, submitted in November 2025) shows that initializing and resetting environments at varied points within the task horizon produces training batches with more temporal diversity and reduces the nonstationarity that synchronized rollouts introduce. The paper also finds that staggered resets scale better as the number of parallel environments grows, compared with naive synchronized rollouts. Resetting an entire fleet of simulator instances all at once produces correlated starting states instead of independent, identically distributed rollouts. If you stagger reset times across the fleet instead, you restore the independence the policy gradient algorithm assumes it has.

Making staggered resets work in practice depends on session-layer isolation. Each training container gets its own session key, and reset(session_id) has to be an atomic, scoped operation that clears exactly that session's state without touching any other session's episode in progress. If you don't enforce session isolation at the simulator layer, staggering the reset schedule doesn't solve anything, because one container's reset can still corrupt another container's mid-episode state through shared mutable data underneath.

Scaling horizontally requires the simulator to run as multiple independent instances with no shared mutable state between them, each instance owning its own session namespace while an orchestration layer routes each container's calls to its assigned instance. A locally runnable, stateful replica of the third-party API, one that remembers state across calls and can be reset by session scope without restarting the server, is the only practical way to recover the Markov property and run episodic training across thousands of parallel containers without their observations colliding. The reset contract exists for this reason: training an agent against a reward signal that means what it claims to mean requires a precondition, not merely a convenience for test writers.

Sources

  1. [2511.21011] Staggered Environment Resets Improve Massively Parallel On-Policy Reinforcement Learning

More in RL Environment Design