Skip to content
markpaper

knowledge/hl/rate-limits.md

vregistry-c914171 · 52.3 KB

Download file
# Hyperliquid — request limits: REST weight, address limits, 429, throttling, egress IP, WS

A reference for developers of bots and tools on HL: how to calculate load, avoid 429, retry without doubling a position, and decide when a second IP is needed. Figures come from HL documentation and measurements; measurements include dates.

## TL;DR

1. HL has **two independent families of limits**:
   - **IP-based:** REST ≈ **1200 weight/min** in total across `/info` and `/exchange`, plus WebSocket limits.
   - **Address-based** (per user address): a buffer of **10,000 requests + 1 request for every $1 of volume**. Once exhausted, 1 request per 10 s remains.

   A second egress IP expands only IP limits; a separate subaccount expands only address limits.
2. `/info` weights:
   - **2:** `l2Book`, `allMids`, `clearinghouseState`, `spotClearinghouseState`;
   - **60:** `userRole`;
   - **20:** almost everything else (`meta`, `userFills`, `openOrders`, `frontendOpenOrders`, `candleSnapshot`, `maxBuilderFee`, `userRateLimit`, etc.). History exports (`userFills*`, `historicalOrders`, and others) and `candleSnapshot` have a documented response-size surcharge (§2.1).

   An exchange request has IP weight **1 + floor(n/40)** for a batch of n orders; under the address limit, each order in the batch counts separately.
3. HL enforces its IP limit over a window **materially shorter than one minute**. In one measurement, a 240-weight burst received 429 even though average usage was well below the minute budget. Therefore a minute counter alone is insufficient: configure sustained rate, burst, and concurrency explicitly and verify them under your own workload.
4. **Every** REST request from one IP (all subsystems and all processes on the machine) must pass through **one** limiter:
   - weight-based token bucket with explicit sustained rate and burst;
   - priorities: urgent CLOSE > orders > background reads;
   - concurrency semaphore;
   - retry with backoff and jitter.

   Two independent limiters on one IP exceed the limit in aggregate.
5. **Retry policy depends on the error class:**
   - **429:** rejected before execution; retry is always safe.
   - **5xx, disconnect, timeout:** unknown outcome. Retry only idempotent operations: info, reduceOnly close, `updateLeverage`.
   - Retry OPEN/INCREASE, TP/SL, and `usdSend`/`sendAsset` transfers **only on 429**.
6. **Concurrent bursts**, not total volume, consume the budget. Typical bursts: a stampede at a cache TTL boundary, synchronously expiring caches, mass REST fallback when WS becomes stale, or a batch of orders emitted at once. Remedies: single-flight, stale-while-revalidate, TTL jitter, warming, and reading state once per tick. A sequential request series is safe.
7. Spreading traffic across multiple **egress IPs** (“lanes”) genuinely multiplies IP limits. But separate buckets are valid **only** for genuinely different source addresses. Check the address at startup: incorrect binding in Node/undici raises no error and silently uses the default IP.
8. WS user tracking is counted **per IP across all connections**. Measurement in 2026-07: ~20 addresses per IP; docs say 10. New connections do not add slots, and moving a subscription to another connection consumes a slot for roughly 60 s.
9. During a 429 storm, the application's own queue stretches execution to **minutes**. Therefore:
   - check the decision age before sending OPEN/INCREASE;
   - use `AbortSignal.timeout` on every fetch to HL.
10. Read the address budget through `userRateLimit` and purchase more through `reserveRequestWeight` (0.0005 USDC per request). For a bot that frequently places and replaces orders at low volume, this is the bottleneck rather than the IP limit.

---

## 9. Address limit: reading, forecasting, purchasing

### 9.1 Mechanics

- A buffer of **10,000** requests + **1 per $1** of cumulative volume (`cumVlm`), verified live on 2026-09-14. Example: a new address has `nRequestsCap = 10000` at zero volume; after $10 of volume, `10010`.
- Each placement = 1, a batch of N = N, `modify` = placement. With frequent replacements it is cheaper to cancel and place again: cancellations use a separate, larger limit `min(limit+100000, limit*2)`.
- Once exhausted: 1 request per 10 s.
- **The limit is shared per address:** multiple bot instances on one account see one remainder. **A subaccount is a separate address** with its own 10,000 buffer. Recommendation: one subaccount per bot.
- **HL has no rolling placement window:** placement is constrained by the address volume-based limit. Calculate an action's cost (number of placements) before starting and do not spend budget on unnecessary replacements.

### 9.2 Reading: `userRateLimit` (weight 20)

```ts
const rl = await hlLimiter.info(W.heavy,
  () => info.userRateLimit({ user: '0xYOUR_ADDRESS' }), 'mon:userRateLimit');
// { cumVlm: string, nRequestsUsed: number, nRequestsCap: number, nRequestsSurplus?: number }
```

`nRequestsSurplus` is capacity above the base cap purchased through `reserveRequestWeight`. Poll roughly every 60 s (never more often than once per 5 s).

### 9.3 Purchasing: `reserveRequestWeight`

Price: **0.0005 USDC per request**; 2,000 = $1, 10,000 = $5 (docs “Reserve additional actions”).

```ts
await hlLimiter.exchange(() => exchange.reserveRequestWeight({ weight: 2000 }),
  { idempotent: false, label: 'trade:reserveRequestWeight' });
// immediately reread userRateLimit and log cap/surplus before and after
```

- **Master account only.** The action does not accept `vaultAddress`; weight is credited to the master account, so do not purchase it for a subaccount.
- **Double counting.** If after purchase the exchange raises both `nRequestsCap` by the chunk size (≥0.9 × chunk) and `surplus`, do not add surplus to the remainder a second time. Otherwise the remainder is overstated and the bot discovers exhaustion only after reaching the limit. The exact field changes after purchase were not verified live (§14).

### 9.4 Budget formulas

```ts
const REQUEST_PRICE_USD = 0.0005; // purchase price, USDC per request (§9.3)

// remaining capacity, including local placements since the last read
function budget(rl: { nRequestsUsed: number; nRequestsCap: number; nRequestsSurplus?: number },
                placesSinceRead: number, surplusCounted = true) {
  const cap = rl.nRequestsCap + (surplusCounted ? Math.max(0, rl.nRequestsSurplus ?? 0) : 0);
  const remaining = Math.max(0, cap - rl.nRequestsUsed - placesSinceRead);
  return { cap, remaining, remainingPct: (100 * remaining) / cap };
}

// exhaustion forecast: each $1 of volume returns +1 request
function hoursToExhaustion(remaining: number, placesPerHour: number, volumePerHourUsd: number) {
  const net = placesPerHour - volumePerHourUsd;
  return net <= 0 || !Number.isFinite(remaining) ? null : remaining / net;
}
```

The remainder grows only through volume and purchase; less frequent placement merely slows consumption. Log the actual placement rate and a “budget lasts ~N h” forecast so exhaustion is not a surprise.

---

## 10. Egress-IP lanes

### 10.1 Why this works

Both IP limits (REST weight ~1200/min and WS user tracking ~20 addresses) are counted **per source IP**. The limit is not per server or account.

Evidence: while the main process held its WS subscriptions through one egress address, a separate probe tracked 21 addresses simultaneously through a second address on the same server, approximately doubling the number of live addresses compared with one IP. Tracking rejections appeared on connections whose local accounting showed 0 and 8 subscriptions.

### 10.2 Design

- **Lanes by purpose:**
  - *monitoring* — background reads, analytics, backfills;
  - *trading* — decisions, stops, orders.
- **Derive the lane from the request label:** `trade:*` → trading, everything else → monitoring. It is the same label used for counters, so telemetry and routing cannot diverge.
- **Orders always use the trading lane:** bind the SDK transport globally to its address.
- **One-off heavy backfills** (dozens of `userFillsByTime`) use monitoring and run strictly sequentially. Monitoring degrades softly; decisions and orders must not compete with background reads for budget.
- **Empty lane address = separation disabled** (default route + one bucket).
- **CRITICAL: separate buckets only for genuinely distinct addresses.** Two buckets on one physical IP emit twice the rate into one limit, guaranteeing a 429 storm. If addresses are unset or equal, every lane shares one bucket.
- **Startup self-check.** In Node/undici, incorrect `localAddress` binding does not error; the request silently uses the OS default IP. A common mistake is nested `new Agent({ connect: { localAddress } })`; only top-level `new Agent({ localAddress })` works. The wrong form creates a 429 storm. At startup, query the external IP through every dispatcher and compare with expectations. If `verified=false`, disable lane separation and restart.
- **A bot isolated on its own IP** is the simplest way not to share the limit with other processes.

```ts
import { Agent } from 'undici';
import * as hl from '@nktkas/hyperliquid';

// TOP-LEVEL option only. undici SILENTLY ignores new Agent({ connect: { localAddress } }):
// the call option is spread after ...options and overwrites it with null; the request uses the OS default IP.
const tradingDispatcher = new Agent({ localAddress: '<TRADING_IP>' });
const monitoringDispatcher = new Agent({ localAddress: '<MONITORING_IP>' });

const tradingTransport = new hl.HttpTransport({ fetchOptions: { dispatcher: tradingDispatcher } as Record<string, unknown> });
const exchange = new hl.ExchangeClient({ wallet, transport: tradingTransport });

// raw fetch for the monitoring lane—with its own dispatcher and timeout
await fetch('https://api.hyperliquid.xyz/info', {
  method: 'POST',
  headers: { 'content-type': 'application/json' },
  body: JSON.stringify({ type: 'meta' }),
  signal: AbortSignal.timeout(20_000),
  dispatcher: monitoringDispatcher,
} as RequestInit & { dispatcher?: unknown });
```

A startup self-check is mandatory: make one request through each dispatcher to an external echo service and compare the actual IP. Check the SDK path (`HttpTransport` with `fetchOptions.dispatcher`) separately from raw `fetch`: it is not confirmed that the dispatcher reaches the real transport request (details in sdk-and-api.md §9 and §12, pitfalls.md §10.1). If the IP does not match, disable lane separation and have every lane share one bucket.

**When another IP is needed.** Exceeding the IP budget is not a ban; it is growing queue latency. A separate IP is justified when bulk polling regularly overlaps trading (§12); set burst/rate parameters explicitly instead of copying them from another workload profile.

---

## 11. WebSocket limits in practice

- **User tracking is per IP, not per connection** (measurement on 2026-07-17 invalidated the “15 per connection” model).
  - New connections from the same IP **do not add slots**: sharding within one IP does not increase capacity.
  - Documented “10 unique users across user-specific subscriptions” is a conservative lower bound for the same limit.
  - The error arrives on the WS error channel: `Cannot track more than 15 total users`. Pushes above the limit silently do not arrive.
- **Measurement pitfall.** HL sends one snapshot at subscription time, so untracked addresses also appear “fresh” immediately afterward. Count only addresses that send a **second** push and wait several minutes. Honest measurements: on a fresh IP, 30 requested → 21 tracked, 9 rejected; on a loaded IP, fewer were live (~18). Therefore the practical estimate is ~20/IP.
- **Ghost slots.** After reconnect or unsubscribe on a *different* connection, HL retains the old connection's tracked users for ~60 s. `unsubscribe` on the **same** connection releases the slot immediately (verified). Two rules follow:
  - **Do not migrate subscriptions between connections:** each migration consumes a slot from the common pool and starts a self-sustaining storm (rejection → no snapshot → stale → resubscribe with migration → new ghost → neighbor rejected). Even below the demand ceiling, only a small fraction of addresses remain fresh.
  - **Resubscribe on the same connection with exponential backoff** (×2 per unsuccessful attempt, cap ~10 min). “Quiet” addresses (an empty account receives only a subscription snapshot) resubscribe forever without backoff.
- **Shard headroom.** A shard filled to the cap rejects resubscriptions, snapshots become chronically stale, all reads fall back to REST at once, and burst 429 results. Keep shard capacity below the cap with headroom (for example, 12 at a cap of 15).
- **Log the WS error channel.** Otherwise a subscription rejection is invisible and data uses REST fallback for hours, accumulating weight.
- **Priority under slot scarcity:** give WS slots first to accounts with open positions (stops need pushes); other addresses use REST fallback and degrade softly.
- **Close WS cleanly at process exit** (stop every connection). Otherwise frequent restarts accumulate hanging connections against the connection limit.
- `allDexsClearinghouseState` does not provide spot balances; read them separately (`spotClearinghouseState`).
- **WS snapshot sanity:** if positions exist but `totalMarginUsed == 0` (payload lacks margin), do not trust the snapshot; use REST.

---

## 12. Pacing bulk requests and backfills

- **The sustainable ceiling for sequential bulk info calls from one IP is ~200–300 requests/min.** A short-burst measurement (1100/min succeeded) is misleading.
- While bulk polling runs at that ceiling, **every other request from the same IP receives 429**.
- **Do not run bulk polling alongside a trading bot on the same IP:** trading also receives 429. Use launch discipline or a separate egress IP.

Empirically selected pauses:

| Scenario | Pause |
|---|---|
| sequential bulk info requests | hundreds of ms between calls to stay around ~200–300 requests/min |
| `userFillsByTime` backfill | 250 ms between pages |
| `candleSnapshot` chunks | 80 ms |

- Run **heavy-request series strictly sequentially**, not as fan-out: the limiter normalizes average rate, while parallel fan-out hits the semaphore and creates a burst.
- **Prefer API dumps.** For fill history from your own builder code, the public S3 dump (`stats-data.hyperliquid.xyz`) costs one GET per day and does not consume the info IP limit.

---

## 13. Pitfalls

| # | What breaks | Why | Correct approach |
|---|---|---|---|
| 1 | Monitoring receives 429 and a close order stalls (order burst into a saturated window without retry) | monitoring has its own post-hoc backoff and trading has another bucket; together they exceed the IP limit | one limiter per IP for every subsystem, info and exchange; change weights and limits in one place |
| 2 | Chronic 429 despite average weight well below budget | the exchange limit window is shorter than a minute; configured burst releases too much at once | reduce explicitly configured burst and concurrency according to live 429 observations |
| 3 | `meta` stampede every 5 min, growing linearly with address count | `getMeta()` on every WS message, TTL without single-flight | single-flight for every heavy-endpoint TTL cache; one reference-data service per process |
| 4 | TTL burst after restart; orders wait in the queue | caches for many keys were created by one restart and expire together | shared SWR cache + in-flight dedup + TTL jitter ±20% |
| 5 | 429 storm from REST fallbacks | WS shard at capacity → stale snapshots → every read falls back to REST together; cache invalidation on every fill | slot headroom, resubscription backoff, 10 s REST floor not reset by invalidation |
| 6 | Trades are silently skipped | 429 on state read → `null` → skip | retry info; `null` ≠ “no positions” |
| 7 | A position increase minutes later reopens a stop-closed position | limiter queue during a storm | decision-age gate for OPEN/INCREASE; CLOSE always proceeds |
| 8 | CLOSE waits behind a batch of INCREASE operations | one high tier | urgent tier for reduceOnly/protective orders; semaphore also transfers slots by priority |
| 9 | Urgent CLOSE hangs for minutes | a stalled socket without timeout occupies a FIFO semaphore slot until undici defaults | `AbortSignal.timeout` on every fetch |
| 10 | Double entry after 5xx | retrying a non-idempotent order after an ambiguous error | OPEN/INCREASE/TP-SL/transfers—retry only on 429; reconcile on 5xx |
| 11 | 5xx was never retried; 502 aborted the tick | error lacks `.status`, classifier checks only the field | propagate `status`; parse `HTTP (\d{3})` from text |
| 12 | Closed-position PnL was not read | info reads sit in the normal queue behind orders and execute after closing | take accounting data from the pre-order snapshot |
| 13 | 429 storm after “IP separation” | two buckets, but traffic actually uses one IP (undici silently ignores `new Agent({ connect: { localAddress } })`) | `new Agent({ localAddress })`; separate buckets only for distinct addresses verified through echo |
| 14 | 1440 weight/min from a simple bot | `frontendOpenOrders` (20) + `l2Book` (2) + `clearinghouseState` (2) every second; `frontendOpenOrders` alone every second is 1200, the full budget | WS `orderUpdates` / `userFills` + REST reconciliation every 10 s |
| 15 | Every process on the IP receives 429 for hours | bulk polling runs alongside the bot | do not overlap; or use a separate egress IP; ~200–300 requests/min is sustainable |
| 16 | Limiter waits forever without an error | a request with weight > capacity can never accumulate enough tokens | reject immediately with “chunk the request” |
| 17 | Tokens are “minted” or withheld | refill uses `Date.now()`, and NTP jumps | `performance.now()` |
| 18 | Entire budget is consumed and other positions are unmanaged | immediate retry of a persistent rejection (margin, lot) | no immediate retry on rejected; cap streak per key |
| 19 | Duplicate close orders in the queue | a slow retry accumulated with a new tick | in-flight Set of keys; re-enqueue does not reset attempts |
| 20 | Health check restarts a healthy process | local throttling counted as an execution error | separate throttling; still expose it as a degradation reason |
| 21 | Orders are not sent for hours while health is green | actions not performed because of budget appear nowhere | separate degradation reason for budget rejections |
| 22 | Replacement limit is exhausted within hours | `modify` and every order in a batch count toward the address limit; replacements are too frequent | replace orders less often, purchase budget, separate subaccount |
| 23 | Address-limit remainder is overstated | `surplus` is double-counted after purchase | if `nRequestsCap` grew by the chunk, do not add `surplus` |
| 24 | WS snapshots do not arrive for hours | tracking rejection is visible only on the error channel; sharding within an IP; migrations create ghost slots | log error channel; do not migrate; resubscribe with backoff; second IP |
| 25 | Startup 429 burst | after restart, the bot places orders and reads state for every account at once | weight-based limiter from the first request; warm caches with a delay after boot |
| 26 | Hanging WS connections accumulate | process exits without closing sockets | stop every connection in the shutdown hook |
| 27 | Raw `fetch` to HL from an auxiliary module | bypasses the common bucket and invisibly consumes IP budget | prohibit raw fetch to HL outside the limiter; analytics uses the monitoring lane |

---

## 14. Open questions / not verified

- **WS connections per IP: 10 or 100?** Earlier note: “~1000 subscriptions and 100 connections per IP”; later: “current official per-IP cap is 10 connections” (medium confidence). Plan for 10 and recheck the docs.
- **User tracking:** docs say 10; measurements say ~20 (18–30); the error says “15 total users.” The adopted model is “~20 per IP in aggregate” (measurement on 2026-07-17 invalidated the “15 per connection” model from 2026-06-11). The exact number probably depends on IP load; retain a safe estimate and priorities.
- **Length of the short IP-limit window** is unknown. It is known only to be materially shorter than one minute: a 240 burst received 429; 80 did not.
- **Sustainable ~200–300 requests/min for bulk polling** versus theoretical 600/min for weight-2 requests. Possible reasons: background load from other processes on the same IP or the short window. Unresolved.
- **An HL concurrency limit as a separate concept** is undocumented; “24 bad, 10 good” is empirical.
- **Weights.** `openOrders`, `frontendOpenOrders`, `extraAgents`, `portfolio`, and `userNonFundingLedgerUpdates` are often treated as 2; according to the docs list (2026-09-13), they are 20. This document adopts the docs list. `orderStatus` = 2 and `allMids` = 2 each come from one source.
- **Response-size weight.** `userFills` and `candleSnapshot` are often treated as fixed 20 regardless of response size. According to the documentation, weight grows with returned item count (§2.1), but the actual surcharge was not measured on our IP; backfill budget figures in §2.3 and §12 use base weight and are understated.
- **`userFillsByTime` history depth** (≤2,000 per response, only the address's ~10,000 latest fills) comes from the documentation and was not verified live.
- **`nRequestsSurplus` after `reserveRequestWeight`:** whether `nRequestsCap`, `surplus`, or both increase was not confirmed live. Log cap/surplus before and after purchase.
- **Exact error text when the address limit is exhausted** and how to distinguish that 429 from an IP 429 in the body were not recorded.
- **Exchange 429 means “guaranteed not executed”** is a working assumption: no duplicates were observed after retrying 429, but documentation did not confirm it.
- **Numeric WS limits** for outgoing messages/min and simultaneous in-flight post requests were not recorded; accounting for exchange actions sent through WS post was not verified.
- **Polling position/order reads over 5+ dexes (HIP-3):** budget +20 per dex × frequency; no live measurements beyond “+45 weight/min for xyz in `frontendOpenOrders`.”
- **Built-in retry/timeout in SDK** `@nktkas/hyperliquid` 0.27.x was not examined: all examples above use the application limiter. Verify that transport does not duplicate retries.

---

Knowledge snapshot: 2026-09; dates of individual checks are in the text. The HL API changes—recheck limits and response shapes.

---

<!-- license-footer -->
_© markpaper authors. Licensed under [CC BY 4.0](LICENSE.md): when publishing or adapting this material, credit “markpaper — Hyperliquid knowledge base” and link to the original and the license._

## 1. Limit map

| Limit | Accounting key | Value | When exceeded | Expanded by |
|---|---|---|---|---|
| REST weight | IP | ≈1200 weight/min, aggregated across `/info` + `/exchange`. In practice there is also a shorter-than-minute window | HTTP 429 | second egress IP |
| Heavy-request concurrency | IP (observation) | a burst of simultaneous heavy requests receives 429: 10 parallel HTTP calls were fine, 24 were not | HTTP 429 | — |
| Address action limit | address (master account or subaccount) | 10,000 + 1 for every $1 of cumulative volume (`cumVlm`). Placement = 1, a batch of N orders = N, `modify` = placement | 1 request per 10 s; the bot effectively stalls | volume, `reserveRequestWeight`, separate subaccount |
| Address cancellation limit | address | separate and larger: `min(limit + 100000, limit * 2)` | — | — |
| WS connections | IP | later note: “current official cap is 10”; earlier note: 100 (see open questions) | — | IP |
| WS subscriptions | IP | ~1000 | — | IP |
| WS user tracking (unique addresses across all user-specific subscriptions) | IP, aggregated across all connections | docs: 10; measurement on 2026-07-17: ~20 (18–30 across measurements) | error channel receives `Cannot track more than 15 total users`; pushes beyond the limit silently do not arrive | new IP only |
| Outgoing WS messages/min | IP | a limit exists; figure not recorded | — | IP |

Special cases:

- The IP limit includes exchange requests (orders) by raw request count. Orders also have an address limit. Therefore orders must consume the common IP budget together with reads, or an order burst drowns reads and itself.
- `stats-data.hyperliquid.xyz` is an S3 bucket, not the info API. The info IP limit does not apply; these requests do not need the limiter.

---

## 2. REST request weights

### 2.1 `/info`

Reconciled with docs `rate-limits-and-user-limits` (read on 2026-09-13).

| Request | Weight | Note |
|---|---|---|
| `l2Book` | 2 | |
| `allMids` | 2 | |
| `clearinghouseState` | 2 | |
| `spotClearinghouseState` | 2 | |
| `orderStatus` | 2 | weight 2 comes from one source; recheck |
| `userRole` | **60** | |
| `meta`, `metaAndAssetCtxs`, `perpDexs` | 20 | loading `meta` + `metaAndAssetCtxs` = 40 per dex |
| `userFills`, `userFillsByTime` | 20 + response-size surcharge | base weight; additional charge for returned fill count, see below |
| `openOrders` | 20 | a common mistake is treating it as 2 |
| `frontendOpenOrders` | 20 | a common mistake is treating it as 2 |
| `candleSnapshot` | 20 + response-size surcharge | one request per coin covers the full window; cost is linear in coin count; additional charge for candle count, see below |
| `maxBuilderFee`, `referral`, `userFees`, `extraAgents` | 20 | `extraAgents` is also often counted as 2 |
| `userNonFundingLedgerUpdates`, `portfolio` | 20 | often counted as 2 |
| `userRateLimit` | 20 | |
| other documented info requests | 20 | |

If your limiter undercounts weight (`openOrders`/`frontendOpenOrders` = 2 instead of 20), the bucket understates usage by as much as 10×. If only your own orders are needed, `openOrders` is no cheaper than `frontendOpenOrders`: both cost 20. Savings come from WS `orderUpdates`, not endpoint selection.

**Response-size-dependent weight** (from the public HL documentation; not measured—reconcile with the current rate-limits page):

- `userFills`, `userFillsByTime`, `historicalOrders`, `userTwapSliceFills`, `recentTrades`, `fundingHistory`, `userFunding`, `userNonFundingLedgerUpdates`, `twapHistory`: base weight 20 plus additional weight for every ~20 returned items.
- `candleSnapshot`: base 20 plus an additional charge for every ~60 returned candles.
- Consequence: a `userFillsByTime` page with 2,000 fills and a `candleSnapshot` response with 5,000 candles cost materially more than 20. Budget a backfill of hundreds of pages by actual item count, not request count. The limiter either reserves conservatively up front or charges additional tokens after the response.

**Fill-history depth** (according to the documentation, not verified): `userFillsByTime` returns at most 2,000 fills per response and only from the address's ~10,000 newest fills. For an HFT address, older history is unavailable through the API; pagination does not recover it. Long history for an active account requires your own recording (WS `userFills`) or the daily builder-fill dump (only for your builder code, §12). Details in fills-and-history.md and backtest-and-data.md.

How to test on your IP: HL does not report remaining IP budget, while `nRequestsUsed` from `userRateLimit` is an address counter and excludes info requests. Test empirically on an idle IP: compare a series of `userFillsByTime` calls returning full 2,000-row pages with a series returning ~10 items, and observe how many consecutive calls precede 429.

### 2.2 `/exchange`

| What | IP accounting | Address accounting |
|---|---|---|
| single `order` | 1 | 1 |
| batch of n orders | `1 + floor(n / 40)` | n |
| TP/SL pair in one call | 1 | 2 |
| `modify` | 1 | same as placement |
| `cancel` | 1 | from the separate cancellation limit |
| `updateLeverage`, switching dex abstraction | 1 | action |

Batching saves IP weight but **not** the address limit.

A useful technique is reserving 2 weight for an exchange action while recording the actual 1 in telemetry, so the dashboard shows actual load rather than reservation.

### 2.3 Cost of typical operations

| Operation | Load |
|---|---|
| Position entry: `updateLeverage` + `order` | 2 exchange actions (1 if leverage is already set): actual IP weight 1–2; accounting with a reservation of 2 per action gives 2–4 |
| Position close | ~2–6 weight |
| `frontendOpenOrders` once per second | **1200/min—the entire IP budget**; together with `l2Book` and `clearinghouseState` once per second: **1440/min, above the 1200 limit** |
| Candles: 50 coins | ≈1000 weight ≈ one minute of budget (at base weight; more with the candle-count surcharge, §2.1) |
| `meta` stampede at the TTL boundary (N concurrent cache misses) | 20 × N: at N = 50, a 1000-weight burst |
| `maxBuilderFee` fan-out to N addresses without a cache | 20 × N in one burst: at N = 50, 1000, nearly the full minute budget |

---

## 3. The short window and burst-class 429

Symptoms of this class:

- HL 429s occur chronically, sometimes in storms;
- average outgoing weight remains well below the configured budget and 1200 limit;
- the cause is bursts: a bucket with 240-weight capacity (~14 s of the minute budget) releases an instantaneous burst that already receives 429;
- retries add pressure, the queue saturates (tokens=0), and orders take minutes.

Fixes:

- reduce the explicitly configured bucket capacity until short-window 429 bursts disappear;
- limit explicitly configured HTTP concurrency: in one measurement 24 simultaneous heavy requests caused 429 while 10 did not; this is an observation, not a universal default;
- separately, maintain slot headroom in WS shards and a floor on REST-fallback frequency (§7, §11).

More observations from the same class:

- **Heavy reads without pauses.** A sequential series of heavy requests without pauses (for example, `userFills` for main + xyz, weight 20, across several accounts) drains a 240 bucket in seconds, and 429s cluster despite low average usage and an empty order queue: the limit is real, while the bot created the burst. Fix: pauses between heavy requests, no parallelism.
- **Synchronous cache expiry.** Caches for many keys created by one restart with the same TTL expire in the same second and produce simultaneous 429s despite low average weight. On a server with low latency to HL, bursts are tighter than on a distant host and these spikes are more pronounced.
- **A 240 bucket was tolerable only under sparse traffic** in this measurement. Do not carry a specific capacity to another workload as a default.

Rule: the limiter and exchange limit absorb a sequential series of heavy requests (`for…of` + `await`). A concurrent burst (`Promise.all` over N requests, a batch of WS messages at a TTL boundary) is dangerous because it emits N × weight at once. Optimize bursts first. There is no need to reduce total volume when it is a small fraction of the limit.

---

## 4. Architecture of an application limiter

### 4.1 Three mechanisms

1. A **weight-based token bucket** smooths average rate under the IP limit. When saturated, callers **wait in a queue** rather than receiving 429.
2. A **concurrency semaphore** limits simultaneous sockets and heavy-read bursts. Fan-out must not open hundreds of connections.
3. **Retry with exponential backoff and jitter** is a safety net. Every attempt acquires tokens again, so backoff respects the limit. Jitter spreads a burst of N parallel calls over time and is the primary relief against IP 429.

The `weightPerMinute`, `burstCapacity`, and `maxConcurrent` parameters are mandatory application configuration. Choose them from the current HL limit, all traffic on that IP, and live 429 observations; the library supplies no hidden values. Request weight and protocol ceilings remain exchange facts, while headroom, burst, concurrency, and retry profile are caller policy.

Priority order:

- **urgent** — reduceOnly CLOSE and protective orders;
- **high** — other orders and reads on the critical path of a trading decision;
- **normal** — background reads and analytics.

Why urgent is a separate tier: otherwise a CLOSE that arrives after a batch of INCREASE operations in the high queue waits for all of them and exits late. Delaying a background read is cheaper than delaying a close.

### 4.2 Implementation (TypeScript)

```ts
type Priority = 'urgent' | 'high' | 'normal';
const TIERS: Priority[] = ['urgent', 'high', 'normal'];
interface Waiter { weight: number; resolve: () => void }

export class TokenBucket {
  private tokens: number;
  private last = performance.now(); // monotonic clock: an NTP jump must neither mint nor withhold tokens
  private q: Record<Priority, Waiter[]> = { urgent: [], high: [], normal: [] };
  private timer: ReturnType<typeof setTimeout> | null = null;

  constructor(private ratePerSec: number, private capacity: number) {
    this.tokens = capacity;
  }

  private refill() {
    const now = performance.now();
    this.tokens = Math.min(this.capacity, this.tokens + ((now - this.last) / 1000) * this.ratePerSec);
    this.last = now;
  }

  private head(): Waiter[] | null {
    for (const t of TIERS) if (this.q[t].length) return this.q[t];
    return null;
  }

  private drain() {
    this.refill();
    let q = this.head();
    // strict tier order; FIFO within a tier
    while (q && this.tokens >= q[0].weight) {
      const w = q.shift()!;
      this.tokens -= w.weight;
      w.resolve();
      q = this.head();
    }
    if (q && !this.timer) {
      const need = q[0].weight - this.tokens;
      const waitMs = Math.max(10, Math.ceil((need / this.ratePerSec) * 1000));
      this.timer = setTimeout(() => { this.timer = null; this.drain(); }, waitMs);
    }
  }

  acquire(weight: number, priority: Priority = 'normal'): Promise<void> {
    // weight > capacity can never accumulate (tokens are clamped to capacity) -> infinite wait without an error
    if (weight > this.capacity) {
      return Promise.reject(new Error(
        `throttle: request weight ${weight} exceeds the bucket capacity ${this.capacity} (chunk the request)`));
    }
    return new Promise((resolve) => { this.q[priority].push({ weight, resolve }); this.drain(); });
  }

  stats() {
    return {
      tokens: Math.floor(this.tokens), capacity: this.capacity,
      urgentQueued: this.q.urgent.length, highQueued: this.q.high.length, normalQueued: this.q.normal.length,
    };
  }
}

export class Semaphore {
  private active = 0;
  private q: Record<Priority, Array<() => void>> = { urgent: [], high: [], normal: [] };
  constructor(private max: number) {}
  get inFlight() { return this.active; }

  async run<T>(fn: () => Promise<T>, priority: Priority = 'normal'): Promise<T> {
    if (this.active >= this.max) {
      await new Promise<void>((r) => this.q[priority].push(r)); // the slot is inherited; do not change active
    } else {
      this.active++;
    }
    try {
      return await fn();
    } finally {
      // priority slot transfer: one FIFO would defeat bucket priority (CLOSE waits behind admitted reads)
      const next = this.q.urgent.shift() ?? this.q.high.shift() ?? this.q.normal.shift();
      if (next) next(); else this.active--;
    }
  }
}

// ---- error classification ----
function statusOf(e: any): number | undefined {
  const s = e?.status ?? e?.response?.status ?? e?.statusCode;
  if (typeof s === 'number') return s;
  // plain Error such as "HL info HTTP 502 for {...}" without .status—otherwise 5xx is never retried
  const m = /\bhttp (\d{3})\b/i.exec(String(e?.message ?? ''));
  return m ? Number(m[1]) : undefined;
}

export function isRateLimitError(e: any): boolean {
  if (statusOf(e) === 429) return true;
  const msg = String(e?.message ?? e).toLowerCase();
  return /\b429\b/.test(msg) || msg.includes('too many requests') || msg.includes('rate limit');
}

const TRANSIENT = ['fetch failed', 'econnreset', 'etimedout', 'socket hang up', 'econnrefused',
  'network', 'timeout', 'eai_again', 'and retry'];

export function isTransientError(e: any): boolean {
  if (isRateLimitError(e)) return true;
  const s = statusOf(e);
  if (s !== undefined && s >= 500 && s <= 599) return true;
  for (let c = e, i = 0; c && i < 6; c = c.cause, i++) if (c instanceof SyntaxError) return true; // malformed JSON
  const msg = String(e?.message ?? e).toLowerCase();
  return TRANSIENT.some((x) => msg.includes(x)); // TimeoutError from AbortSignal.timeout matches 'timeout'
}

// ---- retry ----
const MAX_RETRIES = 4, RETRY_BASE_MS = 250, RETRY_FACTOR = 2.5, RETRY_JITTER = 0.4; // ≈250/625/1560/3900 ms
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));

async function withRetry<T>(attempt: () => Promise<T>, shouldRetry: (e: unknown) => boolean): Promise<T> {
  for (let i = 0; ; i++) {
    try { return await attempt(); }
    catch (e) {
      if (i >= MAX_RETRIES || !shouldRetry(e)) throw e;
      await sleep(Math.round(RETRY_BASE_MS * RETRY_FACTOR ** i * (1 + RETRY_JITTER * (Math.random() * 2 - 1))));
    }
  }
}

// ---- facade: ONE per egress IP ----
export const W = { light: 2, heavy: 20, role: 60 } as const;
const EXCHANGE_WEIGHT = 2; // actual IP weight 1 (+1 per 40 batch orders); reserve conservatively

const bucket = new TokenBucket(config.weightPerSecond, config.burstCapacity);
const sem = new Semaphore(config.maxConcurrent);
const counts = new Map<string, number>();
const count = (label: string) => counts.set(label, (counts.get(label) ?? 0) + 1);

export const hlLimiter = {
  info<T>(weight: number, fn: () => Promise<T>, label: string, priority: Priority = 'normal') {
    return withRetry(async () => {
      await bucket.acquire(weight, priority);
      count(label); // every actual attempt, including retries, is a real request to HL
      return sem.run(fn, priority);
    }, isTransientError);
  },
  exchange<T>(fn: () => Promise<T>, opts: { idempotent: boolean; urgent?: boolean; label: string }) {
    const p: Priority = opts.urgent ? 'urgent' : 'high';
    return withRetry(async () => {
      await bucket.acquire(EXCHANGE_WEIGHT, p);
      count(opts.label);
      return sem.run(fn, p);
    }, opts.idempotent ? isTransientError : isRateLimitError);
  },
  drainCounts() { const out = [...counts].sort((a, b) => b[1] - a[1]); counts.clear(); return out; },
  stats() { return { ...bucket.stats(), httpInFlight: sem.inFlight }; },
};
```

Usage with SDK `@nktkas/hyperliquid` 0.27.x:

```ts
import * as hl from '@nktkas/hyperliquid';

const transport = new hl.HttpTransport();
const info = new hl.InfoClient({ transport });
const exchange = new hl.ExchangeClient({ wallet, transport }); // wallet: viem account for the agent (0xAGENT_ADDRESS)

const state = await hlLimiter.info(W.light,
  () => info.clearinghouseState({ user: '0xYOUR_ADDRESS' }), 'mon:clearinghouseState');

const fills = await hlLimiter.info(W.heavy,
  () => info.userFills({ user: '0xYOUR_ADDRESS' }), 'trade:userFills', 'high'); // critical path of a trading decision

// retry on 5xx/network only for reduceOnly; OPEN/INCREASE only on 429
const reduceOnly = orderParams.orders.every((o) => o.r);
const res = await hlLimiter.exchange(() => exchange.order(orderParams),
  { idempotent: reduceOnly, urgent: reduceOnly, label: reduceOnly ? 'trade:order:close' : 'trade:order:open' });
```

### 4.3 Observability

- **Counters by label.** Every request receives a label; every N minutes, write a breakdown and reset counters. Label scheme:
  - `mon:<type>[:<dex>]` — background monitoring;
  - `trade:<type>` — trading path;
  - `hl:meta`, `hl:meta:xyz` — shared reference data;
  - separate prefixes for analytics and backfills.

  Example line (synthetic): `5min total=104 (~21/min) | mon:clearinghouseState=20 mon:spotClearinghouseState=20 mon:clearinghouseState:xyz=20 mon:userFills=12 mon:userFills:xyz=12 mon:meta=10 mon:meta:xyz=10`. It exposes a `meta` stampede: 10 + 10 requests in the window instead of 1 + 1, one per polled address.
- **The primary duplicate signal** is identical **data** from two subsystems. `mon:meta` + `trade:meta` is a duplicate (global meta is singular). `clearinghouseState` for different addresses from two subsystems is not. Likewise, WS subscriptions for different address sets are not duplicates.
- **Saturation health log.**
  - Variant A logs every 30 s only during saturation: `highQueued > 0`, or `normalQueued > 5`, or `httpInFlight` at the semaphore ceiling.
  - Variant B logs every 60 s (`setInterval(...).unref()`) when there is traffic: `N HL req/60s | queued+Δ tokens=T/C inflight=H | top 6 labels`.
  - A **growing high queue or growing `queued`** means orders are constrained by the IP budget: remove bursts or move traffic to a second egress IP (§10).
- **`stats()` semantics with multiple lanes:** show the bucket with the largest normal queue (observe saturation on the worst lane), and sum `httpInFlight`. `totalQueued` is how many times a caller actually waited for tokens.
- **Deterministic bucket test:** `new TokenBucket(25, 1)` (1 token per 40 ms).
  - First `await bucket.acquire(1, 'normal')` to consume the initial token.
  - Then enqueue `n1, h1, u1, h2, n2`; expected resolution order: `u1, h1, h2, n1, n2` (compare via `JSON.stringify`; set `process.exitCode = 1` on failure).
  - Second case: 4 × high, then urgent CLOSE after 5 ms; CLOSE must resolve first.

---

## 5. Error classes and retry policy

### 5.1 Response classes

| Class | How to recognize it | Was the request executed? | Action |
|---|---|---|---|
| **429 on `/info`** | HTTP 429 / `too many requests` / `rate limit` | no | retry through the limiter with backoff; immediate repetition is pointless |
| **429 on `/exchange`** | SDK `HttpRequestError` with status 429 | **no** (rejected before execution) | safe retry through the limiter; with frequent placements, pause placement for ~10 s |
| **5xx** | status 500–599 or `HTTP 5xx` in text | **unknown** | info—retry; exchange—retry only idempotent operations, otherwise unknown outcome → reconcile (reread orders and position) |
| **network / timeout** | `fetch failed`, `econnreset`, `etimedout`, `socket hang up`, `econnrefused`, `network`, `timeout`, `eai_again`, `TimeoutError`; batch transport timeout = `transport: timeout` | unknown | same as 5xx |
| **malformed JSON** | `SyntaxError` in the `.cause` chain | no effect for info | retry info |
| **exchange response “…and retry”** | substring `and retry` | no | transient; retry |
| **other 4xx** | status 4xx ≠ 429 | no | do not retry |
| **`{status:'err', response: msg}`** | exchange response body | no; the whole batch was rejected | do not retry as 429; preserve the error text |

Errors must carry a status. If raw `fetch` receives `!res.ok`, throw an `Error` with a `status` field, for example, `Hyperliquid API error: ${status}` + `err.status = status`. Otherwise the classifier cannot distinguish a limit from a failure. For example, a 502 from `frontendOpenOrders` aborts the entire tick on the first attempt if `Error("HL info HTTP 502 for {...}")` lacks `.status`, even though timeouts are retried by string match.

### 5.2 What can be retried

| Action | 429 | 5xx | Network / timeout | Why |
|---|---|---|---|---|
| any `/info` | yes | yes | yes | idempotent |
| reduceOnly CLOSE | yes | yes | yes | repetition cannot increase the position (`idempotent: params.reduceOnly`) |
| `updateLeverage`, switching dex abstraction | yes | yes | yes | idempotent |
| OPEN / INCREASE order | yes | **no** | **no** | the order may have filled; repetition doubles entry |
| `positionTpsl` (TP/SL pair) | yes | **no** | **no** | duplicate placement after ambiguous 5xx is worse than missing protection |
| `usdSend` / `sendAsset` | yes | **no** | **no** | the transfer may have succeeded; retry with a new nonce doubles it. HL deduplicates only an identical nonce, so “do not transfer twice” belongs to the caller's ledger/state machine |
| `reserveRequestWeight` | yes | no | no | costs money |

### 5.3 Backoff profiles

| Context | Attempts | Delays | After exhaustion |
|---|---|---|---|
| Trading-bot limiter (trading + monitoring) | 4 | `250 × 2.5ⁿ × (1 ± 0.4·rand)` ≈ 250 / 625 / 1560 / 3900 ms, total up to ~6.3 s | throw |
| Offline history backfill (backtest) | 8 | start 1500 ms, ×2, cap 30,000 ms; wait and retry on 429 | error with final status and response body |
| Lightweight REST bot | 3 | 429 → 400 ms ×2; other non-ok → 400 ms without doubling; network → ×2 | `undefined` → `ok:false` → **skip tick** |
| Bot with frequent placements, exchange 429 | — | pause placements for ~10 s | — |

Four attempts on the trading path are justified as follows:

- a late CLOSE is better than no close;
- extra seconds are not critical for info;
- under load, completing an OPEN is better than losing the trade. But stale decisions are filtered by the freshness gate in §8.

```ts
// Minimal info request with 429 retry (Node 20+, no dependencies)
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));
async function info(body: unknown, tries = 3): Promise<any | null> {
  for (let i = 0; i < tries; i++) {
    try {
      const r = await fetch('https://api.hyperliquid.xyz/info', {
        method: 'POST',
        headers: { 'content-type': 'application/json' },
        body: JSON.stringify(body),
        signal: AbortSignal.timeout(20_000),
      });
      if (r.status === 429) { await sleep(3000); continue; }
      if (!r.ok) { await sleep(1200); continue; }
      return await r.json();
    } catch { await sleep(1200); }
  }
  return null; // null = failure/degradation; an empty exchange response looks like [] or {}
}
```

### 5.4 What to do with “no data”

- **Lightweight bot** (all state is read on every tick): do nothing on failure; skip the tick. Do not place orders using partially read state.
- **Bot where state reading is a trade precondition:** 429 while reading state → `null` → the trade is **silently skipped**, with no order or error. Info reads are idempotent: retry the read instead of skipping the trade.
- **Do not cache errors.** `null` means an HL error and is not cached, so the next call retries. An empty array is valid state and is cached. Do not write `null` over the last known value in a write-through cache.

---

## 6. Your throttle ≠ the exchange limit

- **Different statuses.** Using one `SKIPPED` status for two cases (held by your limiter / actual exchange rejection due to a limit or minimum) hid real exchange complaints and produced “alert/recovery” pairs. Keep separate statuses, counters, and logs.
- **Throttling is not failure.** Count placements held by your limiter separately from execution errors, or a health check restarts a healthy process. Rejection by local policy (lot, minimum order) is not an execution error either.
- **But it cannot be silent.** An action not performed because of budget must surface as a separate degradation reason, or orders can remain unsent for hours under green health.
- **Count only meaningful deferred actions.** Budget rejections—yes; routine deferred actions (for example, lot minimum) can occur hundreds of times per hour, making health permanently red and as useless as permanently green.
- **An IP-weight bucket does not cover the address limit** on order count. Backoff and a separate address-budget module do (§9).
- **“Never 429 by construction”** is true only for average load. The exchange's short window still produces 429 on spikes, so retry is required even with a perfect bucket.
- **Bootstrap after restart** (all addresses at once: clearinghouse + spot + xyz + meta) produces a one-time 429 spike absorbed by limiter and retry. In steady state, 429 should be 0.

---

## 7. Weight-saving patterns

| Pattern | Essence | Purpose / effect |
|---|---|---|
| **Single-flight before every TTL cache for a heavy endpoint** | cache stores an in-flight Promise; every parallel caller reuses it | without it, requests at a TTL boundary equal the number of concurrent callers. If `getMeta()` runs on every WS message, 50 active addresses create ~50 fetches = a ~1000-weight burst every 5 min |
| **One process-wide service for global reference data** | one `meta`/`meta:xyz`, `allMids` cache per process for every subsystem | two subsystems with independent caches for the same `meta` double requests; after consolidation there is exactly 1 + 1 request (main + xyz) per TTL window. Memoize the derived `Map<coin, …>` by universe-array identity |
| **Single-flight key includes priority** | `address:priority` | a high request must not inherit a background normal request's queue position for the same address. Two concurrent `userFills` calls (main + xyz) without deduplication = 40 weight per address |
| **Stale-while-revalidate** | return stale immediately, update in the background, deduplicate in-flight | caches for a heavy request (for example, `maxBuilderFee`, weight 20) across many keys, created by one restart with equal TTL, expire together: an N × 20 burst drains the bucket and orders wait in the queue for tens of seconds |
| **TTL jitter ±20%** | `ttl = round(base × (0.8 + random × 0.4))`, once at write time | spreads expiry across N keys |
| **Warm at startup** | warm heavy-data SWR caches ~20 s after boot | otherwise the first post-restart request waits for a cold burst across every key |
| **Short cache + single-flight for mids/book** | `allMids` (2) and `l2Book` (2): 2 s TTL, separate in-flight per dex, cache errors | N concurrent calls on a cold cache produced N × 2 weight, competing with orders |
| **Read state once per tick** | read an account's orders, positions, and equity once per tick and reuse them across all logic; one `allMids` per dex | reads do not repeat for every consumer; next step is WS `allDexsClearinghouseState` snapshot instead of REST |
| **WS first, REST as degradation** | state from a WS snapshot = 0 REST; REST only when snapshot is stale (for example, >90 s) | position read chain: WS → 30 s cache → **10 s full REST per account** → single-flight → REST |
| **REST-fallback frequency floor that invalidation cannot reset** | separate snapshot of the latest REST response, untouched by cache invalidation | otherwise with stale WS, each own fill invalidates the cache and a frequent sweep hits fresh REST on every tick: a `clearinghouseState` avalanche feeds the 429 storm |
| **Write-through cache instead of a separate fetch** | store balance where it is already read (UI render, periodic sweep); do not add a dedicated balance fetch | do not write `null`; a stale value means “unknown” |
| **Short cache for a repeated heavy request** | 30 s cache for a `userFills` response | a repeated request for the same address immediately after the first does not pay twice |
| **WS for orders and fills** | WS `orderUpdates`, `userFills` + REST reconciliation every 10 s | instead of `frontendOpenOrders` every second (1200/min, the entire budget; with `l2Book` + `clearinghouseState` each second: 1440/min) |

---

## 8. Timeouts and freshness during a storm

- **`AbortSignal.timeout(20_000)` on every raw `fetch` to HL**; 60 s for large background exports.
  - Without a timeout, a stalled socket occupies a semaphore slot until undici defaults (~minutes).
  - If the semaphore is FIFO without priorities (priority lives only in the bucket), 10 stalled info requests block urgent CLOSE too.
  - The classifier matches `TimeoutError` on `timeout`, so info retries work.
- **Timeouts on the critical path** through `Promise.race`, with an `.unref()` timer. Do not raise them beyond tens of seconds: they retain critical-path resources.
- **Decision freshness gate.**
  - Why: during a 429 storm, decision-path reads (free stables, mid, meta) sit in a saturated limiter for minutes, and an increase can execute minutes later and reopen a position already closed by a stop.
  - Rule: immediately before sending, compute `Date.now() − decidedAt` (from the moment of decision). OPEN/INCREASE older than the threshold → `SKIP stale_decision_age:Ns`. CLOSE always proceeds. Threshold 0 disables the gate.
- **Take accounting data (PnL before close) from the pre-order snapshot:** a read after the order in a saturated queue may arrive after closing and return `null`.
- Use **high priority for reads on which a fresh trading decision or position close depends.** Background recalculation is normal.
- **In-flight guard against accumulating retries.** If a close retry from the previous tick is still running (waiting on a mutex or limiter under 429), skip the tick for that key. Re-enqueueing does not reset `attempts`/`firstQueuedAt`. Likewise, in-flight flags for polling and reconciliation prevent a loop from overlapping itself.
- **Immediate retry on a persistent rejection = hot loop.**
  - “Retry immediately” both aborts the remainder of the tick and cancels the loop delay. For a rejection that cannot clear in 20 ms (insufficient margin, size not expressible as a lot), this becomes a zero-delay loop that burns the entire budget while other positions go unmanaged.
  - Rule: no immediate retry after a rejected IoC. Cap immediate retries per `account:coin` key, then `log.error` and resume the normal loop.

---
All files