General pitfalls in developing trading bots on HL that do not fit the topics of orders, WS, limits, or accounts: number precision in JS, “not read” versus “empty,” replica lag, races, caches, remembered state, closes, restarts, duplicate processes, networking, and observability. Every rule includes the failure mechanism and how to prevent it.
TL;DR
JSON.parsesilently rounds identifiers above 2^53. Cancellation by the rounded id succeeds—the exchange answers OK—while the order remains in the book (not yet observed on HL; see §1.1). Store order identity only as a string. The checkString(Number(x)) === String(x)proves nothing:xwas already rounded during parsing.- “Not read” ≠ “empty.” A failed or partial read (of even one dex) must produce
ok:false/complete:falseand skip the tick. - One HL read is not a fact. An info replica lags by seconds. Your own order appears in
clearinghouseStateafter ~0.5 s. A WS frame and the HTTP placement response arrive out of order. Use a debounce (2 consecutive reads), a positions → orders → positions “fence,” and trust your own writes. - Serialize processing per key with a promise-chain lock. The key is the coin without side: HL has one net position per coin. There must be no
awaitbetween checking a flag (stop, lifecycle) and acting. Schedule ticks withsetTimeoutafter completion, notsetInterval. - Caches. Cold or expiring meta caches need single-flight, or N × 20 weight is spent in one burst. Event-driven invalidation needs a lower bound on read frequency.
- A remembered number remembers its formula, subject, and generation. Store a peak with a formula fingerprint and a subject (account) or generation key.
- Closing is the critical path. No gate may block CLOSE. It needs retries and urgent priority. Close from actual exchange state, not a local registry. A feature switch must not disable management of positions already open.
- One bot = one wallet. A reconciler that treats every account order as its own destroys a neighboring bot's book, and the neighbor retaliates in kind.
- Networking.
undicisilently ignoresnew Agent({ connect: { localAddress } }). Every HTTP request to HL needs a hard timeout. Detect 429 with\b429\b, not a substring. - Instrumentation lies silently.
armed=trueand/health: okcan be false. A flag must represent real operation, and every “deferred / skipped” path must appear in health.
1. Numbers, precision, strings
1.1 Identifiers beyond 2^53
Number.MAX_SAFE_INTEGER = 2^53 − 1 = 9007199254740991 ≈ 9.007e15.
On HL, oid arrives as a number and is 64-bit in the protocol. Precision loss has not been observed yet, but the trap is potentially the same: monitor oid growth.
How it breaks (observed on an exchange with ids around 1e16 > 2^53, where the API returns the id as both string and number: the string is exact while the number is already corrupted by JSON.parse). Cancellation using the rounded numeric id succeeds: the transaction is valid and the exchange answers OK, but the order remains live.
- Such orders become uncancellable: every tick “cancels” them again, wasting requests.
- Two nearby ids round to the same number. The reconciler treats two orders as one: the second is neither matched nor cancelled, which looks like “duplicates” even though the cause is id precision.
Rules:
- The identity of an order (or any new identifier field near 1e16) is a string. A number is acceptable only as an internal handle with a “string ↔ number” table. Send the id to the exchange as a string.
- Check precision loss before parsing by comparing with a string twin.
String(Number(x)) === String(x)proves nothing. - A nonce represented as a JS
numberhas the same trap.
// Detect precision loss by comparing the number with its string twin from the same response
function assertIdExact(idStr: string, idNum: number): void {
if (!Number.isSafeInteger(idNum) || BigInt(idStr) !== BigInt(idNum)) {
throw new Error(`order id precision loss: ${idStr} parsed as ${idNum}`);
}
}
1.2 Sizes, prices, ticks
| Pitfall | Why it breaks | Correct approach |
|---|---|---|
Math.floor(sz * 10 ** szDecimals) | 0.29 * 100 = 28.999999999999996, so floor returns 28 instead of 29 | Math.floor(sz * 10 ** d + 1e-9) or string arithmetic |
| Prices calculated arithmetically with a step | Binary error accumulates with a step such as 0.05 | Number(v.toFixed(10)); compare prices with tolerance a >= b − 1e-9 |
| Comparing a live order with the target as floats | Target 41.237 is sent as 41.24, comes back as 41.24, and never equals the numeric target. The result is perpetual cancel + place on every tick | The comparison key is the quantized price string, exactly what is sent in the order's p field. Pass the live order through the same function |
| Inferring the tick from decimal places in a price string | Formatters strip trailing zeroes (String(p) on HL). Price 40 → "40" → “tick 1” → a ±$0.50 tolerance band around an actual 0.001 tick; at price 2 the band is ±25%. This can match a target to another order at another price | An order matches a target when both prices produce the same string on the exchange grid. Percentage tolerance is only a fallback, with no “floor” |
| A “ticks” parameter near a power of ten | Just above 100, the tick is 10 times larger than just below. An “N tick” offset tuned below the boundary moves the order 10 times farther above it | Do not specify an offset as a fixed number of ticks: the tick changes tenfold at a power-of-ten boundary |
| Size rounds down to the lot near the $10 minimum | If order size near the $10 floor rounds down and is not calculated from the order price, notional becomes slightly below $10 (for example, $9.92); every tick is rejected, and retries over tens of minutes exhaust the address request limit | Calculate size from the order price and round up to the lot until it reaches $10 (with margin). Do not retry a persistent rejection (minimum, margin) every tick—apply a cooldown |
| Numbers in info responses | sz, px, szi, accountValue, limitPx, closedPnl, startPosition are strings | Number(...); compare with zero using a 1e-9 threshold |
Math.max(...ts) over a timestamp array | Works for thousands of elements; reaches the argument limit with tens of thousands | reduce |
1.3 ?? and NaNs that silently disable logic
Number(u.maxLeverage ?? 0). A missing field becomes0, notundefined, andMath.min(leverage, maxLeverage)silently yields leverage 0 → size zero. Cap leverage only whenmaxLeverage > 0. HL currently always returnsmaxLeverage > 0; this guards against a schema change or new market types.Number('{...}')= NaN, andav < NaN * kis alwaysfalse. If JSON is written into the peak's kv value instead of a number, the protection trigger switches off without a single log line. Rules:- store the number as a numeric string and metadata (provenance) in a separate label;
- compare “peak matches label” as strings:
Number()loses low-order digits; - separate label segments with
~, not+:+occurs in exponential notation (1e+21).
1.4 Time, string keys, names
- ISO without milliseconds. When API data is written alongside archive rows that have second precision, format it byte-for-byte the same:
new Date(ms).toISOString().replace(/\.\d{3}Z$/, 'Z'). In string comparison,'Z'(0x5A) >'.'(0x2E), so a fill in the same second “moves” into the next period. This is a systematic monetary error. Intl.DateTimeFormat('en-US', { hour12: false })formats midnight as24. To calculate “minutes after midnight” (for example, in ET for market hours), usehourCycle: 'h23'andformatToParts(year/month/day/hour/minute/weekday).- HIP-3 coin names contain a colon (
xyz:CL,xyz:SP500). A composite key${wallet}:${coin}:${side}cannot be parsed withsplit(':'). Read coin and side from the original record. - The main HL perp dex is encoded as the empty string
''.list.filter(Boolean)silently removes it. For example, comparing dex sets can interpret “narrowed to onexyz” as an expansion, and disappearance of the main dex (and part of equity with it) will not reset the remembered peak. - Prefix keys. When cleaning by suffix
:${coin}without a separator,BTCmatchesXBTC; require a colon before the coin.
2. “Not read” ≠ “empty”
2.1 Invariants
- Degraded read → skip the whole tick. A read function returns
ok:false; its caller never treats that as “the account is empty.” Otherwise the bot decides there are no orders and places duplicates, or decides there are no positions and erases state. - Failure to read even one dex (main or HIP-3) → skip the ENTIRE tick for every dex. A partial picture produces wrong budget and exposure decisions.
- A failed
clearinghouseStateread withdex:'xyz'(422 unknown dex, 429, network) returns{ positions: [], accountValue: 0, ok: false }.- Previous
xyz:*positions are treated as still active. Otherwise a read failure looks like disappearing positions and leads to false decisions (state reset, closing). - Swallow an error for an additional dex internally (do not crash the main polling loop), but propagate it as
ok=false.
- Previous
- Open orders across multiple dexes are read with
Promise.allSettledand return{ orders, complete };complete=falseif any dex does not respond.- A negative conclusion on incomplete data (“no orders”) is unreliable.
- A positive conclusion (“an order exists”) remains valid: missing orders cannot invalidate it.
- An empty order list is a common legitimate account state and cannot be distinguished from a failure. Therefore a network error must not become
[], and every such error is logged.
- A zero position size in a snapshot means “no data,” not a close. The exchange sometimes transiently returns a position with zero size. Still add the position key to the active set; skip the position itself with a warning.
- Plausibility checks:
accountValue < totalMarginUsed—such an account would already be liquidated, so the response is corrupt;- do not trust a WS snapshot with open positions but
totalMarginUsed == 0; use REST. OtherwisemarginRatio= 0 and the opening risk cap never fires; - on a failed free-stables read, return 0: conservatively assume there are no free stables.
- Missing account address → a “fake flat” snapshot. An empty position map and synthetic flat
{ position: 0, orders: [], ok: true }are indistinguishable from a confirmed empty account. Never delete records or state from such a snapshot: a live HL position would remain without a DB record or management. - A read failure never counts toward the flat debounce (see §3.3).
2.2 Fail-closed or degrade—the criterion
Ask which is more expensive: acting on a corrupt number or not acting at all?
| Fail-closed (the error costs money) | Degrade instead of throwing (throw blocks protection) |
|---|---|
One malformed universe item invalidates the whole universe: asset ids shift | A spot-read failure must not skip a guard / wind-down tick |
NaN in szi invalidates the whole position map | A foreign positional TP/SL with sz: "0.0" must not crash order reads, or the engine runs cycles without reading |
| A duplicate stable invalidates the entire spot contribution | An xyz-meta failure must not crash the main meta |
Cancellation without explicit success = not cancelled |
Order decoder: validate identity and role flags strictly for every order, but validate px/sz fail-closed only for orders managed by the engine. Foreign orders must decode: they count as exposure. A common mistake is validation order—checking px/sz before the relevance filter that would have discarded the order anyway.
2.3 Last-known-good cache
Rules for equity (and likewise for stables):
- Update the cache only from a complete, trusted, positive read.
- During degradation, return the cache within a 10 min TTL with
equityFresh=false: it cannot authorize exposure growth. - No cache →
ok:false(skip tick). - A complete, honest zero read deletes the cache.
- A hard-invalid number (negative, nonnumeric) is never replaced with a positive cache:
accountValue=0,ok=false.
Without these rules, the sequence 100 → 0 → error returns 100 with ok=true.
An untrusted read (implausible for a dex, incomplete, or with degraded stable parsing) does not enter the smoothing median window and does not become the baseline. The engine returns the previous median, or 0 if no window exists (fail-closed). With stale sizing inputs it suppresses exposure growth while other work continues.
2.4 “The field was parsed but not consumed”
If a meta field (for example, a market-mode flag) is parsed but discarded when converting to internal meta, the executor never sees it: the code looks ready, but behavior does not change. In review, check who consumes every parsed meta field.
3. Snapshot consistency and HL lag
3.1 What is known about lag
- An info replica can lag matching by seconds. A lagging replica may return an empty account order book for one tick.
- After your own order,
clearinghouseState(REST and WS snapshot) updates after ~0.5 s. - The placement response and WS frames (
orderUpdates,userFills) arrive out of event order. A fill or terminal status for anoidcan arrive before the HTTP response that contains thatoid. - REST requests for
clearinghouseStateandopenOrdersare not atomic: fills may occur between them.
3.2 Snapshot “fence”: positions → orders → positions
Positions and open orders are two independent HTTP reads. An order filled between them produces orders=[] plus a stale position: a state that never existed, so decisions based on it (such as wind-down) are false.
Rule: read positions before and after open orders (and spot), and accept the snapshot only when position sizes match. Otherwise set positionsOk=false and equityFresh=false: sizing stays on previous values and positions are not cut because of a phantom zero. Dangerous actions (reset, wind-down) are forbidden on an unstable snapshot.
const INFO = 'https://api.hyperliquid.xyz/info';
async function info<T>(body: object): Promise<T> {
const r = await fetch(INFO, {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify(body),
signal: AbortSignal.timeout(10_000), // see §10.3
});
if (!r.ok) throw Object.assign(new Error(`HTTP ${r.status}`), { status: r.status });
return r.json() as Promise<T>;
}
const sizes = (st: any) =>
new Map<string, number>(st.assetPositions.map((p: any) => [p.position.coin, Number(p.position.szi)]));
function sameSizes(a: Map<string, number>, b: Map<string, number>): boolean {
const keys = new Set([...a.keys(), ...b.keys()]);
for (const k of keys) if (Math.abs((a.get(k) ?? 0) - (b.get(k) ?? 0)) > 1e-9) return false;
return true;
}
const withDex = <T extends object>(body: T, dex: string) => (dex ? { ...body, dex } : body); // the main request has NO dex key
async function stableSnapshot(user: string, dex = '') {
const before = await info<any>(withDex({ type: 'clearinghouseState', user }, dex));
const orders = await info<any[]>(withDex({ type: 'frontendOpenOrders', user }, dex)); // includes reduceOnly/isTrigger/cloid
const after = await info<any>(withDex({ type: 'clearinghouseState', user }, dex));
return { state: after, orders, stable: sameSizes(sizes(before), sizes(after)) };
}
- For the main dex, the
dexkey is omitted entirely from the request body (notdex: '')—details in market-data.md. - Order reconciliation requires
frontendOpenOrders, notopenOrders: only it includesreduceOnly,isTrigger, andcloid—details in orders.md.
3.3 One read is not a fact: debouncing
- Act on “no orders” only after the second consecutive such read: a lagging replica may return an empty order list for one tick (§3.1), and acting on it creates duplicates.
- Act on “position is flat” only after 2 consecutive flat reads. A read failure (
positionsOk=false) never counts as flat. - The debounce-counter key must include the lifecycle row's generation (see §6.3).
3.4 A local book instead of openOrders every second
You cannot poll openOrders every second: it weighs 20 against an IP-wide budget of 1200 per minute for everything. Local order-book design:
| Source | What it provides | Frequency |
|---|---|---|
| Placement response | oid | per event |
| Cancellation response | confirmation | per event |
WS orderUpdates / userFills | fills and cancellations between ticks | stream |
REST openOrders | reconciliation only | ~10 s |
REST clearinghouseState (weight 2) | position | 2 s |
Five rules for reconciling a local book with REST openOrders:
- An order placed after the reconciliation request was sent (
placedAt ≥ sentAt) or younger thangraceMs = 10 sis not considered missing. - An order cancelled by us after the request was sent does not “resurrect.” Use tombstones with a TTL of 60 s: a snapshot requested before cancellation or fill has no right to bring the order back.
- Confirm disappearance with two consecutive reconciliations, not one.
- Remaining size from reconciliation only decreases:
sz = min(local, exchange). - Remember a terminal update or fill for an
oidnot yet recorded (pending, TTL 20 s—longer than the 10 s HTTP timeout), and apply it when the placement response arrives.
Adopt foreign account orders into the local book as “unmanaged.”
Scenario: “the WS order fills before the placement response.” An orderUpdates open frame for an unknown oid is adopted as an unmanaged order. The order then fills or is cancelled (tombstone). Only afterward does the placement HTTP response arrive with the same oid as resting. It must not resurrect: placement checks tombstones and pending terminal updates and treats the order as already gone. Remaining size = min(target − pending.filledSz, pending.openSz, adopted.sz).
Accepting a REST position snapshot (position = latest accepted snapshot + subsequent WS fills):
- accept the snapshot if it matches the estimate within
tol = 1.5 × 1e-9; - or if it equals the estimate without a tail of the k newest fills (the node has not applied them yet). Those fills remain “after the snapshot”;
- otherwise there is a conflict: the position does not change. If the same value arrives twice consecutively, accept it (the exchange is authoritative; WS missed an event);
- while the conflict is unresolved, treat the snapshot as not received. If the conflict persists beyond a configured threshold, pause placement of new orders.
3.5 Expected position after a fill
After a confirmed fill, a snapshot can be internally stable but still lag the write. Reject a snapshot that does not match the expected position as stale.
function positionAfterFill(
before: number, side: 'B' | 'A', fillSize: number, reduceOnly: boolean,
): number | null {
if (!Number.isFinite(before) || !Number.isFinite(fillSize) || fillSize <= 0) return null;
if (reduceOnly) {
if (before > 0 && side === 'A') return Math.max(0, before - fillSize);
if (before < 0 && side === 'B') return Math.min(0, before + fillSize);
return null;
}
return before + (side === 'B' ? fillSize : -fillSize);
}
3.6 Medians and “pairs that never existed”
Suppose a guard peak stores a perp leg and two spot companions (total, free), while each field is taken from its own median over its own window array. The result is a state that never existed: free == total (zero reserve) with a nonzero perp leg, even though the reserve must mirror it. While the perp leg transitions 0 → X, most ticks in the window do not yet reflect the reserve, so the free median comes from another state. A protection trigger compares an invented value with a real one, sees zero growth, and shuts down an account where nothing happened.
Rules:
- A remembered pair or triple of values must have been observed at one moment. Choose one complete observed tick (for example, nearest to the accepted perp-leg peak, preferring ticks where the reserve mirrors the perp leg) and take every field from it. The median of each field independently is mathematically sound and physically impossible.
- Keep window arrays aligned by index. Write an unread value as
nullinstead of skipping the record. Conditional writes silently desynchronize arrays, andspot[i]stops describingperp[i]. - Check identities before a monetary decision (reserve = perp leg; capital = perps + free spot). Per-field aggregation breaks them, making this a ready detector.
4. Races and concurrency
4.1 Race catalog
| Race | Symptom | Remedy |
|---|---|---|
| Two decisions for one coin less than 500 ms apart (open and rapid increase) | clearinghouseState has not updated; both decisions see “no position” → double entry | Mutex per (account, coin) |
| The same decision is processed twice (for example, a trigger arrives from both polling and WS), milliseconds apart | A mutex serializes but does not make the snapshot fresh → the position is opened twice | positionExists = snapshotHasPosition || ownRecord != null. Write your own position record synchronously on FILLED OPEN. The second decision sees the position → SKIP |
| Buy and sell orders for the same coin from different bot components on one account | Both see “no position” and send non-reduceOnly. HL has one net position per coin → netting becomes a mess and two records represent one position | Mutex key ${account}:${coin} without side |
Ticks overlap under setInterval | A tick longer than the interval starts another on top → duplicate orders and cancellation races with HL error Order was never placed… | Let a tick schedule itself with setTimeout after completion. Duplicates self-heal on the next clean cycle |
An await between checking “stopped” and acting | The stop commits within the window while the same tick sends a market IoC → a leveraged position appears seconds after stop; pausing on the next tick does not undo the fill | Put every network operation (listing check, meta) above the gate. Code after the gate is synchronous |
| A state request outlives an account change | A response for account A writes into caches (last-good stables, etc.) after A was cleared and B now uses the same identifier | After await, verify that the address has not changed and discard the response before writing caches. Same for agent-key association checks |
| Semaphore: release then reacquire | The gap between active-- and waking a waiter lets a new run() jump the queue → active > max, weakening 429 burst protection | Transfer the slot atomically (snippet below) |
| Finalization after a lock-free phase | While alive / orderStatus was checked without a lock, the position closed and reopened with a new record and new TP/SL pair. Finalizing against the new record would orphan the live position | Under the lock, reread the record and compare identity: openedAt and SL oid and TP oid |
4.2 Promise-chain lock (per key)
const locks = new Map<string, Promise<void>>();
async function withKeyLock<T>(key: string, fn: () => Promise<T>): Promise<T> {
const prev = locks.get(key) ?? Promise.resolve();
const run = prev.then(fn);
const tail = run.then(() => undefined, () => undefined); // rejection does not break the chain
locks.set(key, tail);
try {
return await run;
} finally {
// delete only if the tail is our task; compare the SAME reference stored in the Map
if (locks.get(key) === tail) locks.delete(key);
}
}
// execution: withKeyLock(`${account}:${coin}`, ...) — without side
Pitfall: if next.catch(...) is stored in set, but finally compares it with a new next.catch(...), these are different promises. === is always false, cleanup never runs, and the Map grows forever (leak by unique keys).
4.3 Per-account execution lock
Live order batches and control commands (pause, stop) share one per-account queue. The queue stores a tail and pending counter; it is deleted from the Map when pending === 0. Why: pause or stop must not land between cancelling and replacing orders, or a position is left without protective TP/SL.
4.4 Semaphore with atomic slot transfer
class Semaphore {
private active = 0;
private queue: Array<() => void> = [];
constructor(private max: number) {}
async run<T>(fn: () => Promise<T>): Promise<T> {
if (this.active >= this.max) {
await new Promise<void>((resolve) => this.queue.push(resolve));
// woken by slot transfer—the slot is already ours; do not change active
} else {
this.active++;
}
try {
return await fn();
} finally {
const next = this.queue.shift();
if (next) next(); // transfer the slot without decrementing
else this.active--;
}
}
}
// Invariants: active++ only when active < max; the queue is nonempty only when active == max.
Priorities exist only in the token bucket. The semaphore is plain FIFO: if its slots are occupied by stalled requests, an urgent order also waits. The only slot protection is a timeout on every fetch (§10.3).
4.5 A loop without overlapping ticks
let running = true;
let pendingTrigger = false;
let wake: (() => void) | null = null;
function triggerTick() { pendingTrigger = true; wake?.(); } // WS trigger (fill, etc.)
function sleepUntilNextTick(ms: number): Promise<void> {
if (pendingTrigger || !running) return Promise.resolve();
return new Promise((resolve) => {
const t = setTimeout(() => { wake = null; resolve(); }, ms);
wake = () => { clearTimeout(t); wake = null; resolve(); };
});
}
async function loop(reconcileMs: number) {
while (running) {
pendingTrigger = false; // clear BEFORE the tick: a fill during it starts the next one
try { await tick(); } catch (e) { console.error(e); }
await sleepUntilNextTick(reconcileMs);
}
}
4.6 Isolation and bulk operations
- Bulk operational closing uses
Promise.allSettled, notPromise.all. One task throwing synchronously would reject the entirePromise.allalong with results from neighbors that already closed, and the report would show “total failure” despite partial success. A rejected task counts asclosed:false. - Distinguish
null(“service disabled; nothing ran”) from[](“nothing to close”). Otherwise an emergency close with a disabled service reports “closed 0 of 0.”
5. Caches: stampedes and freezes
5.1 Meta stampede
- Symptom. During the meta TTL window (5 minutes), instead of 1 + 1 requests (main meta and xyz meta), the system makes N, where N is the number of simultaneous
getMeta()calls: N × weight 20 instead of 20. A similar pattern occurs withuserFillsat startup. - Cause. N concurrent calls (for example, a
Promise.allacross several reads, orgetMeta()on every incoming WS message) all miss a cold or expiring cache. Without in-flight deduplication, every one sends a request for the same global meta. The stampede grows linearly with N: 50 simultaneous calls are ~50 × 20 = 1000 weight in one burst every 5 minutes against a 1200-per-minute limit. Duplicate caches across subsystems are only the tip of the problem.
Remedy: one process-wide cache and single-flight.
type Meta = unknown;
let metaCache: { value: Meta; at: number } | null = null;
let metaInFlight: Promise<Meta> | null = null;
async function getMeta(fetchMeta: () => Promise<Meta>, ttlMs = 5 * 60_000): Promise<Meta> {
if (metaCache && Date.now() - metaCache.at < ttlMs) return metaCache.value;
metaInFlight ??= fetchMeta()
.then((value) => { metaCache = { value, at: Date.now() }; return value; })
.finally(() => { metaInFlight = null; });
return metaInFlight;
}
5.2 Cache-pitfall catalog
| Pitfall | What happens | Correct approach |
|---|---|---|
| Startup / restart = burst | Initial reads for every account start together: clearinghouse requests + meta + userFills. Throttle and retry absorb the burst, but frequent restarts multiply it. Exclude the startup window from measurements as unrepresentative | Avoid unnecessary restarts; stagger startup |
| Event-driven invalidation bypasses the cache | Position cache with short TTL: every own FILLED order (dozens during a fill storm) invalidates it, and frequent stop sweeps hit REST every tick → a REST avalanche under 429 | Lower bound on repeated REST reads (10 s) that survives invalidate. Reset it only when changing client or account |
| Accounts with fresh WS are never polled over REST | WS does not provide spotClearinghouseState. Spot stables (part of accountValue) freeze at process startup | Poll each account over REST at least once every 5 min, even with fresh WS |
| Cold meta cache during close | A close order fails with the internal error “asset meta not found” when there is no retry | Warm meta before sending closes; on a miss, single-flight reload and retry instead of silently skipping. A new xyz listing can also miss |
6. Remembered state: formula, subject, generation
6.1 “A remembered number remembers its formula”
An equity peak in kv is a number from the past, while a trigger divides today's reading by it. This is correct only while both values use the same formula.
- If the peak was written by an incorrect formula (for example, “perps + all spot,” double-counting), it survives the formula fix in kv.
- A correct reading against it shows a false large “drop” → latch and market close.
- The threshold would not have been crossed against a correct peak.
The remedy is a provenance label stored next to the number: formula version, its components (for example, perps + free spot or perps only), and the dex list.
- Raise the label version together with a formula change.
- A label may only reset the peak (disabling the trigger until rebasing) and never create a latch.
- Narrowing the dex list (including
'', §1.4) also resets the peak.
6.2 Keys without a subject
- A remembered-state key without an account identifier is inherited by a new account after an account change, and comparison with the old value creates false changes. Delete such keys when changing accounts.
6.3 Generation in debounce keys
If a flat-streak counter uses ${account}:${coin}, while pause and activate increment the row's generation, the counter survives pause → resume and stop → resume: one flat blip before pause leaves streak=1, and the first blip after resume makes streak=2. The debounce degenerates to one observation, potentially causing a market close. Flat cleanup that treats “no address” as confirmed flat also deletes the operational row.
Correct: key ${account}:${lifecycleGen}:${coin}; “no data” ≠ “confirmed zero.”
6.4 Merging userFills responses by dex
userFillsignores thedexparameter (verified on 2026-06-11; fills-and-history.md): a second request “for xyz” returns the same fills, and concatenating the two responses doubles the array. Anything calculated or calibrated on duplicated data behaves differently after deduplication: remove the extra request while preserving data shape, or recalculate (recalibrate) derived values and thresholds.
7. Closing and exit are the critical path
7.1 Invariants
- Entry gates never block CLOSE. Otherwise a position remains without an exit until SL/TP.
- Do not collapse defense-in-depth barriers on different paths into one: the close handler has several callers, including bulk close.
- Close from the actual exchange position, not a local registry. A bulk close that iterates a local registry closes only positions with records. A position missing from the registry (for example, opened manually or after a failed write) is invisible and can remain open for hours. Correct: iterate accounts and close
(coin, side)fromclearinghouseState. - A close requested before HL exposes a just-filled entry (~500 ms) is queued for completion when there is a fresh own position record. The reconciler rereads state on the next tick and completes reduceOnly to flat. Without an own record, this is a no-op.
- Do not increase a position queued for forced close. For example, during a partially filled stop exit (the latch is set only after full close), before the next reconciliation.
- Delete a local position record only when HL does not see it AND the grace window has elapsed (an invalid
openedAtis stale). Do not delete a fresh record: the next rapid decision on the same asset would bypass the guard and double or reverse the position.
7.2 Where closing breaks
-
Close burst. Bulk-closing many positions at once produces a mix of failures:
429 Too Many Requests, a cold meta-cache miss, or no price inallMids. Without retries, positions remain open for hours. Remedy: retries, urgent throttle priority, request-limit headroom, warmed meta, and a manual force-close script through the same reduceOnly path. -
A feature switch must not disable management.
- If an opening feature's
ENABLED=falsereturns fromstart()immediately, the engine and its management/closing of positions already open disappear too; only the protective SL remains. - The same applies to
DRY_RUN=trueif the dry-run return also exists in the close branch.
Invariant: the switch for an opening feature does not disable management of open positions. Disable new entries separately while the engine continues existing positions through close.
- If an opening feature's
-
Emergency flatten does not take mid from a WS frame: during halt, nothing updates it. During halt/flatten, read mid over REST.
7.3 reduceOnly is exchange semantics, not a constant
On HL, a resting reduceOnly order cannot open a position. Typical places where an engine silently relies on this:
- a terminal close leaves TP orders during closing;
- full close skips the cancellation fence;
- replacement TP orders after cancellation are sized from the pre-batch snapshot.
Test this property explicitly with isolated fixtures: treating resting reduceOnly like a normal limit would incorrectly allow a position to cross zero into the opposite side. Fixtures do not establish real exchange behavior; preserve the distinction between observed semantics and test assumptions.
8. Reconciliation and execution
8.1 IoC and execution markers
-
Race: “a resting order filled between snapshot and IoC.” The positions → orders → positions fence proves no fills occurred during the reads, but not after the second position read and before sending the IoC. The result is double action: an entry fills → stale IoC fills the gap again; TP fills → IoC reduction cuts again (reduceOnly prevents reversal but not excess reduction).
Remedy: when there are resting orders for a coin, cancel them first and reread state (even a successful cancellation could have partially filled before ack), then decide from fresh state whether IoC is required. If the fill already brought the position to target, do not send IoC.
-
Advance an “executed” marker only after actual execution. If an IoC was deferred or omitted from the batch while the marker advanced, execution silently does not happen and no instrument sees it. Count attempts for a specific order from the placement result of that order, not from “batch sent”: a batch may contain only cancellations.
8.2 Remainder, coin set, and position direction
- Treat a remainder that cannot be expressed as an order as exhausted. If it is below the lot,
minSz, or minimum notional ($10 on HL), no order can consume it,remaining <= 0never becomes true, and orders are cancelled and replaced tick after tick without a journal entry. A size that oscillates with equity makes this flapping permanent.
const remainderUnplaceable = (r: number) =>
r <= LOT_EPSILON || floorSz(r) <= 0 || r < minSz - 1e-12 || r * refPx < minNotionalUsd;
- A position without open orders is not flat. If the coin set is built only from coins with open orders, a position without orders becomes invisible: it is treated as flat and market-closed. Build the set from coins with orders or a position.
- Position direction = sign of
szi. Direction inferred from the composition of open orders can be the opposite, causing opposite-side protection to close the correct position.
8.3 Budget, denominators, sizing
-
Endogenous denominator → oscillation. If exposure-cap “capacity” is calculated from values reserved by the engine's own resting orders (for example, free spot moves into hold), a feedback loop results:
- orders reserve spot → capacity falls → the cap reduces positions at market → orders are cancelled → capacity recovers;
- the cap multiplier jumps from tick to tick and positions are cut without an external cause.
Rule: verify that every sizing input is independent of the engine's own orders.
-
A rejection by your own limiter ≠ skip because of the exchange minimum. If both use one status (for example,
SKIPPED), they are indistinguishable. A minimum-size skip is persistent (repeats until position or balance grows); your limiter's rejection is temporary (the window clears in seconds). Use a separatethrottledflag.
8.4 Error streak and pause
If a consecutive-exchange-error counter that triggers pause is reset only by successful placement, pause becomes permanent: there are no placements during pause. The streak must decay by itself after a configured period without new errors.
8.5 “The bot treats every account order as its own”
- A reconciler that cancels every account order not matching its targets treats every account order as its own. Two such bots on one HL wallet endlessly cancel each other's books with real money. Invariant: one bot = one wallet (subaccount). Do not enable live until addresses are reconciled.
- Manual orders placed by the account owner on the same coins are also cancelled.
- The order-ownership marker (a cloid prefix on HL) is part of the contract with the live book. Changing the marker scheme between deployments orphans the entire resting book: orders without the marker are treated as foreign, are neither cancelled nor matched. Change it only with a migration.
9. Processes, restarts, duplicate instances
9.1 Restart and shutdown
- Death between cancel and re-place. The “cancel → place again” cycle must be atomic with respect to process shutdown. A short process-manager shutdown timeout followed by SIGKILL leaves a position without TP/SL. SIGTERM means “wait for the current cycle to finish.”
- Adopt an open position from a previous run through
clearinghouseState; do not close it.
9.2 Duplicate processes
| Situation | What breaks | Rule |
|---|---|---|
| Two trading bots on one wallet | Mutual destruction of books (§8.5) | Separate wallet or subaccount |
| Frequent restarts | Every startup is a REST burst across all accounts (§5.2) | Do not restart unnecessarily |
9.3 Modules with side effects and tests
- Put pure functions (decision rules, backtests) in modules without network or DB side effects. Importing a side-effectful module in a test starts the network or storage.
- A client module that reads configuration (API URL) at import time. If a test overrides env after imports, the client calls the production Hyperliquid API instead of the mock. Import such modules in tests only after overriding env; test pure functions separately.
10. Network, HTTP, WS
10.1 Binding the outbound IP in undici
In lib/core/connect.js, undici takes localAddress from call-level options and spreads it after ...options, always overwriting the value from connect:{} with null. It is populated only by the Client's top-level option. There is no error: traffic silently uses the default IP, so two throttle buckets hit one limit and create a 429 storm.
import { Agent } from 'undici'; // pin "6.21.2" without a caret
const laneA = new Agent({ localAddress: '<EGRESS_IP_A>' }); // works
// new Agent({ connect: { localAddress: '<EGRESS_IP_A>' } }) // SILENTLY ignored
// dispatcher does not pass tsc: with no "lib" in tsconfig, target ES2022 pulls in lib.dom,
// whose RequestInit has no dispatcher. Use a local cast at the boundary.
const res = await fetch('https://api.hyperliquid.xyz/info', {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({ type: 'meta' }),
dispatcher: laneA,
signal: AbortSignal.timeout(10_000),
} as RequestInit & { dispatcher?: unknown });
- An egress self-check is mandatory at startup: one request per lane to an external echo service (not in the hot path), then compare the actual addresses.
- Create separate throttle buckets only for genuinely distinct addresses. HL budgets by IP.
- Undici version. The npm
Agentdrives Node's built-in fetch by duck-typing the handler protocol. Within 6.x this is compatible: pinned 6.21.2 works against Node's built-in 6.24.1 and the egress self-check passes. Do not use 7.x: the handler API was redesigned (onRequestStartinstead ofonConnect/onHeaders) and was not tested with Node 20 fetch.
10.2 Detecting 429
Raw includes('429') matches 429 inside an oid or large number (14290, 0x…429…) and may broaden retries for a non-idempotent OPEN after an ambiguous failure → double entry.
function isRateLimited(err: any): boolean {
const status = err?.status ?? err?.response?.status ?? err?.statusCode;
const msg = String(err?.message ?? err ?? '').toLowerCase();
return status === 429 || /\b429\b/.test(msg) || msg.includes('too many requests') || msg.includes('rate limit');
}
10.3 Timeouts
- Sequential tick + long timeout = freezing every account. If a tick processes accounts sequentially, one stalled HTTP call (undici default ~300 s × retries) freezes order management for all accounts, including guard closes. Use
AbortSignal.timeout(10_000)on every HL request; put retries and backoff in the throttle. Fail fast, let retry handle it. - A 20 s fetch timeout is the only protection for semaphore slots from stalled sockets (§4.4).
10.4 WS: ghost slots and capacity measurement
- Ghost slot for ~60 s. If you unsubscribe on one WS connection and subscribe on another, HL keeps the address occupied for ~60 s and rejects the new subscription. Even below the capacity ceiling, this creates a self-sustaining storm: after several minutes only a small share of addresses remain fresh and the others are rejected. Unsubscribe on the same connection releases the slot immediately.
- The silence threshold must account for positions. An address with no positions can legitimately be silent on
allDexsClearinghouseState. With one short threshold (for example, 90 s), it repeatedly becomes “stale” → resubscription with migration to another shard (ghost slot ~60 s). If resubscription updateslastSeen, backoff never engages because the next tick resets the attempt counter. With many subscriptions for positionless addresses, ghost churn happens simultaneously → pressure on the users/connection cap →Cannot trackand 429. Remedy: long threshold for confirmed-empty addresses (for example, 15 minutes) + exponential backoff. - Measuring WS capacity. Immediately after subscribing, the tracked-address count is inflated because some “fresh” addresses sent only one-time snapshots. An honest method waits several minutes and counts only addresses that send a second push:
- when subscribing 30 addresses, 21 were tracked and 9 rejected (fresh IP);
- fewer were live on a loaded IP (~18);
- practical estimate: ~20 per IP.
- WS degradation moves reads to REST fallback. If the fallback interval is large (minutes), every consumer suffers. Therefore slots for critical feeds take priority over auxiliary subscriptions.
10.5 Weights and frequencies
| Request | Weight | Typical frequency |
|---|---|---|
openOrders | 20 | reconciliation only, ~10 s |
clearinghouseState | 2 | 2 s for your own account |
meta (and HIP-3 dex meta) | 20 | 5 min cache + single-flight |
| Budget | 1200 per minute per IP | for everything |
11. Observability: instruments that lie silently
armed=trueon an inactive guard. If the field means “enabled in config and has a peak,” it remainstruewhile evaluation is never called (positionsOk=false) or exits at the first line (untrusted equity). Correct:armed= the guard is actually evaluating now, plus fieldsevaluating,notEvaluatingReason,lastEvalAgoMs(an age above several ticks means execution never reaches the guard).- Deferred actions must appear in health. If “action deferred” is set on only one path while another branch writes only a log, a position can lag its target for weeks with
state: okand emptydegradedReasons. Same for the lot minimum. Every “deferred / skipped” path must mark a health signal. - Cancellations are not logged → flapping is invisible. Cancelling and replacing orders every few minutes can happen without a single journal line. Log cancellation or at least a per-coin churn counter.
- Do not invent error codes. If an exchange rejection code is unknown, capture it from the journal on first observation.
12. Open questions / not verified
- HL
oidand 2^53. oid is 64-bit in the protocol; the value at which precision loss becomes real was not verified. Storing oid as string or BigInt is already safer. - Median smoothing of snapshots. There are two criteria for selecting a complete snapshot: median by
accountValueor nearest by perp leg to the accepted peak. The principle is the same (the whole snapshot was observed); which criterion is better is not verified. - Undici compatibility. There is a requirement to use “exactly the version built into Node”; pinned 6.21.2 was verified against built-in 6.24.1. Practical rule: compatible within major 6.x; 7.x is not.
- WS capacity. Some measurements show a cap of ~15 users per connection and others ~20 tracked addresses per IP. These are likely different limits (connection and IP), but their relationship is unknown. See the WebSocket topic.
- REST fallback for position reads during WS degradation. Intervals from 15 s to 10 minutes are found in practice; a long interval blinds stops throughout degradation. A safe value is unknown.
Knowledge snapshot: 2026-09; dates of individual checks are in the text. The HL API changes—recheck limits and response shapes.
© markpaper authors. Licensed under CC BY 4.0: when publishing or adapting this material, credit “markpaper — Hyperliquid knowledge base” and link to the original and the license.