Skip to content
markpaper

knowledge/hl/websocket.md

vregistry-c914171 · 50.8 KB

Download file
# Hyperliquid — WebSocket

Reference guide to the WebSocket API of Hyperliquid (HL) for those who are writing trading tools: bots, account monitoring, and market-data clients. Facts are from working with a live socket; where a fact is taken from documentation and not re-verified, it is marked as such.

## TL;DR

1. **URL `wss://api.hyperliquid.xyz/ws`**. Subscribe: `{"method":"subscribe","subscription":{...}}`; unsubscribe: `{"method":"unsubscribe","subscription":{...}}`. Every server frame arrives in an envelope `{channel, data}`. Each subscription is confirmed by a separate frame with `channel:"subscriptionResponse"`.
2. **The server closes the socket after 60 seconds without messages.** Send `{"method":"ping"}` every 30 seconds; the response comes in `channel:"pong"`. If no frames (including pong) are received for 65 seconds, call `terminate()` and reconnect.
3. **Main limit — number of tracked user-addresses per IP.** Documentation mentions 10 unique users in user-subscriptions. A measurement on 2026-07-17 showed around 20 addresses in total across all connections. New connections from the same IP do not get new slots; capacity adds only with a different egress-IP. The overall limit of subscriptions is approximately 1000 per IP.
4. **Limit rejection comes in `channel:"error"` as a string `Cannot track more than 15 total users`**, and there are no data for the extra addresses. If you do not log the `error` channel, the rejection is completely invisible.
5. **Ghost-slots.** If an address is unsubscribed on one connection and resubscribed on another, the slot remains occupied for about 60 seconds. After a reconnect, the old connection also holds its slots for around 60 seconds. Therefore, moving subscriptions between connections is not possible; re-subscription can only be done on the same socket. On the same socket, `unsubscribe` frees up the slot immediately.
6. **Snapshots and updates.** `userFills` first sends a frame `isSnapshot:true`: this is history, not new fills. `allDexsClearinghouseState` sends a snapshot upon subscription (even to rejected addresses), then only pushes changes afterward. A blank account remains silent after the first snapshot. `l2Book` only pushes updates when the book changes, so silence does not mean disconnection.
7. **Reconnect:** backoff 1 second → ×2 → 30 seconds. Wait at least 60 seconds if the close code is `1008` (too many connections from IP). Reset backoff only after stable work with confirmed subscriptions, not on `open`. After a reconnect, resend all subscriptions.
8. **WS gives speed; truth lies in REST.** Positions, orders, and balances should be periodically reconciled via REST (`clearinghouseState`, `openOrders`, `spotClearinghouseState`). Spot balances are not available via WS. Any WS cache must have a REST fallback.
9. **Time race conditions.** `orderUpdates` with status `open` comes before the HTTP response for placing an order. After closing a position, the WS snapshot may still show it for several seconds. The first snapshot after subscribe arrives in 0.5–2 seconds.
10. **WS transport SDK `@nktkas/hyperliquid` 0.27.1 dies completely after 3 unsuccessful reconnects (`maxRetries=3`).** A long-lived bot needs its own client on the `ws` package with an infinite reconnect loop.

---

## 1. Connection and Protocol

| What | Value |
|---|---|
| URL (mainnet) | `wss://api.hyperliquid.xyz/ws` |
| Subscribe | `{"method":"subscribe","subscription":{"type":"<type>", ...params}}` |
| Unsubscribe | `{"method":"unsubscribe","subscription":{...same object...}}` |
| Ping | `{"method":"ping"}` → response `{"channel":"pong"}` |
| Envelope of any frame | `{ channel: string, data?: unknown }` |
| Subscription confirmation | `{ channel:"subscriptionResponse", data:{ method:"subscribe", subscription:{ type, user?, coin? } } }` (one per each subscription) |
| Request error | `{ channel:"error", data:"<string>" }`: rejected subscription, limit, malformed request, invalid address |
| Server idle-timeout | 60 seconds without messages |
| Close code `1008` | policy violation. Actually means exceeding the number of WS-connections from IP |
| Close reason `Expired` | observed before reconnect, after which subscriptions were not restored en masse (see pitfalls) |
All numbers in payload (prices, sizes, `accountValue`, `szi`, `unrealizedPnl` and so on) come as **strings**.

Normalize addresses to lowercase (`user.trim().toLowerCase()`) and validate them with `/^0x[0-9a-f]{40}$/` before sending a subscription. The server may return the address in a different case than it was sent.

### Options for package `ws` (8.x)

```ts
import WebSocket from 'ws';

const ws = new WebSocket('wss://api.hyperliquid.xyz/ws', {
  perMessageDeflate: false, // do not waste CPU/GC on zlib-inflate for each frame in the hot path
  handshakeTimeout: 15_000,
  localAddress,             // egress-IP; undefined = default OS route
});
```
- `perMessageDeflate:false`: if the server is network-wise close to HL, the bandwidth is not narrow, and it's more beneficial not to spend time on decompression in `onMessage`.
- `localAddress` works: `ws` 8.20.1 forwards unknown options through `https.request → createConnection → tls.connect`. Only `socketPath`/`hostname`/`protocol`/`timeout`/`method`/`host`/`path`/`port` are overwritten. Thus, one process can open WS with different egress-IP (see §4).

---

## 2. Subscriptions and Formats

### Summary Table

| `type` | Parameters | `channel` Frames | Behavior | Live Verification |
|---|---|---|---|---|
| `l2Book` | `coin`, `nSigFigs?`, `mantissa?` (aggregation per documentation, not verified) | `l2Book` | frame only on book change | yes |
| `userFills` | `user`, `aggregateByTime?: boolean` | `userFills` | first frame `isSnapshot:true` (history), then new fills | yes |
| `orderUpdates` | `user` | `orderUpdates` | array of order status changes | yes |
| `allDexsClearinghouseState` | `user` | `allDexsClearinghouseState` | full perp state across all dex; snapshot on subscribe, then push on changes | yes, primary source |
| `webData3` | `user` | `webData3` | aggregated blob; single place with account mode string `userState.abstraction` | yes, for one-time reading |
| `webData2` | — | — | snapshot of user state; may lag by seconds after position closure | only mentioned |
| `trades`, `candle`, `bbo`, `activeAssetCtx`, `allMids` | see "Market Subscriptions" below | same-named | forms per public documentation | **no**, verify against live socket |
| `userEvents`, `userFundings` | — | — | not dissected here | **no**, verify against official documentation |
User-specific subscriptions (`userFills`, `orderUpdates`, `allDexsClearinghouseState`, `webData*`, `userEvents` and so on) consume **one common** unique address limit (§4).

### `l2Book`

```json
{"method":"subscribe","subscription":{"type":"l2Book","coin":"BTC"}}
```

The frame comes only when the order book changes (as stated in the official documentation and confirmed practically). On a quiet market, there may be no frames for seconds, which is normal.
- If the WS-order book is older than your freshness threshold, fetch the order book via REST `l2Book` (weight 2). If it's still stale, do not make trading decisions based on this order book.
- Disconnection should be determined by the absence of **all** frames, including `pong`, and not just the lack of `l2Book`.
### `userFills`

```json
{"method":"subscribe","subscription":{"type":"userFills","user":"0xYOUR_ADDRESS"}}
```

`data = { user, fills: [...], isSnapshot?: boolean }`. A fill has at least `tid` and `time`.
- After each (re)subscription, the first message received is a snapshot with `isSnapshot:true`: history of recent fills, including those that occurred while the socket was down. **This is not new events.**
- How to handle snapshots: In PnL ledger, take only fills with `time >= sessionStartTime` that have not yet been seen by `tid`. The position from the snapshot **do not change**, it is restored by REST `clearinghouseState`.
- `aggregateByTime` aggregates partial fills over time as the same-named flag in REST `userFillsByTime` (the parameter is declared, but not verified in practice).
- Filter by `user` in lowercase.
```ts
if (msg.channel === 'userFills') {
  const d = msg.data as { user: string; fills: Array<{ tid: number; time: number }>; isSnapshot?: boolean };
  if (d.user.toLowerCase() !== me) return;
  if (d.isSnapshot) {
    for (const f of d.fills) {
      if (f.time >= sessionStartTime && !seenTid.has(f.tid)) { seenTid.add(f.tid); ledger.add(f); }
    }
    scheduleSync(); // positions and orders only from REST
    return;
  }
  for (const f of d.fills) if (!seenTid.has(f.tid)) { seenTid.add(f.tid); onFill(f); }
}
```
### `orderUpdates`

```json
{"method":"subscribe","subscription":{"type":"orderUpdates","user":"0xYOUR_ADDRESS"}}
```

`data` — array:
```ts
type OrderUpdate = {
  order: {
    coin: string;
    side: 'B' | 'A';
    limitPx: string;   // string
    sz: string;        // REMAINDER, string
    oid: number;
    origSz: string;    // string
    timestamp: number;
  };
  status: string;      // 'open' — in the book; 'badAloPxRejected' — post-only rejection; full list — §13
  statusTimestamp: number; // event time (fallback: order.timestamp)
};
```
- `open` comes **before** the HTTP-response `exchange` on placement. The local order book should be able to "adopt" an unknown oid from WS and then merge it with the placement response. Consider the case when the order is already filled or canceled by the time of the response.
- oids in one batch of placements go consecutively.
- A failure from Alo comes **twice**: in the response on placement (`statuses[i].error` = `Post only order would have immediately matched…`; SDK throws an exception even if neighboring orders stood) and in `orderUpdates` with status `badAloPxRejected`. Without deduplication, the failure counter is exactly doubled.

### `allDexsClearinghouseState`

```json
{"method":"subscribe","subscription":{"type":"allDexsClearinghouseState","user":"0xYOUR_ADDRESS"}}
```

One subscription returns the perp-account state immediately for all dexes, including HIP-3. No REST weight is spent on reading.
```ts
type Payload = {
  user: string;
  clearinghouseStates: Array<[dexName: string | null, inner: {
    marginSummary?: { accountValue?: string; totalMarginUsed?: string };
    crossMarginSummary?: { accountValue?: string };
    spotState?: { totalRawUsd?: string }; // declared in the documentation, not verified live (see Open Questions)
    assetPositions?: Array<{ position?: {
      coin?: string; asset?: number; szi: string; unrealizedPnl: string;
      marginUsed: string; entryPx: string; leverage?: { value: number | string };
    } }>;
    time?: number;
  }]>;
};
```
- `dexName` `''` (or `null`) means the main perp HL, `'xyz'` — HIP-3 dex xyz. There can be other dexes (`flx`, `vntl`, …).
- Each frame is a **authoritative full snapshot**. If a dex is not in the array, it means there is no state on that dex; this is not an error in loading.
- In `position`, `coin` may not be present, only the index `asset`. Then the name is taken from `meta.universe[asset].name` (for xyz — from the xyz meta).
- XYZ coins should be formatted as `xyz:COIN` if there is no `:` in the name.
- Subscription confirmation comes from frame `subscriptionResponse`, data arrives in `channel:"allDexsClearinghouseState"`.

Parsing position and margin (main + xyz, other dexes are skipped):
```ts
if (msg.channel !== 'allDexsClearinghouseState' || !msg.data) return;
const user = msg.data.user?.toLowerCase();
const positions: Position[] = [];
for (const entry of msg.data.clearinghouseStates) {
  if (!Array.isArray(entry) || entry.length < 2) continue;
  const [dexName, inner] = entry;
  const isXyz = dexName === 'xyz';
  if (dexName !== '' && dexName != null && !isXyz) continue; // skip flx/vntl/...
  const accountValue = toNum(inner?.marginSummary?.accountValue);
  const marginUsed = toNum(inner?.marginSummary?.totalMarginUsed);
  for (const ap of inner?.assetPositions ?? []) {
    const p = ap?.position; if (!p?.coin) continue;
    const szi = toNum(p.szi); if (szi === 0) continue;
    let coin = p.coin; if (isXyz && !coin.includes(':')) coin = `xyz:${coin}`;
    positions.push({ coin, side: szi > 0 ? 'LONG' : 'SHORT', size: Math.abs(szi),
      leverage: toNum(p.leverage?.value), unrealizedPnl: toNum(p.unrealizedPnl),
      marginUsed: toNum(p.marginUsed), entryPx: toNum(p.entryPx) });
  }
}
```
Bring the WS result to **the same form** as the REST path (`clearinghouseState` main + xyz). Then both sources go into one handler, and the logic does not diverge.

### `webData3`: Account Mode

The exact string of the collateral mode (DEX abstraction) is **only** in WS `webData3` in the field `data.userState.abstraction`. It is not present in REST (`webData2`, `clearinghouseState`, `userDexAbstraction`).

| `userState.abstraction` | Meaning |
|---|---|
| `"disabled"` | manual: dex are isolated, each has its own collateral |
| `"unifiedAccount"` | new default: common collateral for all dex |
| `"dexAbstractionEnabled"` | legacy abstraction, enabled through app.hyperliquid.xyz |
In the manual-akount, in other fields, you can see `dexAbstractionEnabled:false`, `userDexAbstraction:false`. In `data` there is also `perpDexStates[]` (multi-dex blob).

Short-lived socket reading once (justified only for rare polling, e.g., every 15 minutes):

```ts
import WebSocket from 'ws';

export function readAccountAbstraction(user: string, timeoutMs = 6000): Promise<string | null> {
  return new Promise((resolve) => {
    const ws = new WebSocket('wss://api.hyperliquid.xyz/ws');
    let settled = false;
    const done = (v: string | null) => {
      if (settled) return; settled = true;
      clearTimeout(timer); try { ws.close(); } catch {}
      resolve(v);
    };
    const timer = setTimeout(() => done(null), timeoutMs);
    ws.on('open', () => ws.send(JSON.stringify({
      method: 'subscribe', subscription: { type: 'webData3', user },
    })));
    ws.on('message', (raw) => {
      try {
        const msg = JSON.parse(raw.toString());
        if (msg.channel === 'webData3') done(msg.data?.userState?.abstraction ?? null);
      } catch { /* ignore */ }
    });
    ws.on('error', () => done(null));
  });
}
// await readAccountAbstraction('0xYOUR_ADDRESS') → 'disabled' | 'unifiedAccount' | 'dexAbstractionEnabled'
```
### Market Subscriptions

> According to HL public documentation, not verified in live environment (except noted). Before usage, compare form fields with a live socket and `.d.ts` SDK.

| subscription | `channel` / `data` | Purpose |
|---|---|---|
| `{type:'trades', coin}` | `trades`: array `{ coin, side: 'B'\|'A', px, sz, time, hash, tid, users: [buyer, seller] }` | trade feed: secondly OHLCV, VWAP, execution model |
| `{type:'candle', coin, interval}` | `candle`: `{ t, T, s, i, o, c, h, l, v, n }`, same form as REST `candleSnapshot` | live bars `1m`+; last bar updates to close |
| `{type:'bbo', coin}` | `bbo`: `{ coin, time, bbo: [bid \| null, ask \| null] }`, level `{ px, sz, n }` | cheap top-of-book (spread, mid) without full book |
| `{type:'l2Book', coin, nSigFigs?, mantissa?}` | `l2Book`: `{ coin, time, levels: [bids, asks] }` | order book; `nSigFigs` aggregates levels |
| `{type:'activeAssetCtx', coin}` | `activeAssetCtx`: `{ coin, ctx: { markPx, oraclePx, midPx?, funding, openInterest, dayNtlVlm, dayBaseVlm, prevDayPx, premium?, impactPxs } }`. Keys in live frame BTC 2026-09-23 (read-only): `dayBaseVlm, dayNtlVlm, funding, impactPxs, markPx, midPx, openInterest, oraclePx, premium, prevDayPx` | mark/oracle for stop orders on mark, current funding |
| `{type:'allMids', dex?}` | `allMids`: `{ mids: Record<coin, string> }` | all coins' mid prices in one frame |
Rules (follow from the mechanics described in this file):
- Market subscriptions do not consume the unique user-address limit (§4), but count towards the overall subscription cap on IP (around 1000).
- Trades frames may repeat after a reconnect. Deduplication by `tid` is mandatory, as with `userFills`.
- All numbers come in as strings. `side` in `trades` — aggressor side (taker).
- Silence in `trades` and `l2Book` on a quiet market is normal. Activity is checked by default across all frames, including `pong` (§6).
- For HIP-3, `coin` is passed with a prefix (`xyz:TSLA`), as in REST `l2Book`.

For more details on data recording for backtesting — see backtest-and-data.md.

### What's Missing from WS

- **Spot balances** (analog of REST `spotClearinghouseState`). Stablecoins USDC/USDT/USDH/USDE are obtained via REST. Any equity formula involving spot requires periodic REST requests (with caching for minutes).
---

## 3. Snapshot and Updates: How to Understand Data Freshness

| Channel | What Comes Immediately | What Follows | What Silence Means |
|---|---|---|---|
| `userFills` | `isSnapshot:true` with history | only new fills | normal, no trades |
| `allDexsClearinghouseState` | full snapshot within 0.5–2 seconds after subscribe, **even if the address will not be tracked later** | full snapshot on every change; for open positions push on mark-price movements (unrealizedPnl), usually every few seconds | there are positions and silence over 90+ seconds — subscription has failed. Empty account is silent always, this is normal |
| `l2Book` | current book | frame on every change | normal on a quiet market |
| `orderUpdates` | — | status changes | normal, no orders |

Consequences:
- **The fact of the first message says nothing about the liveness of the subscription.** A live subscription is confirmed only by **repeated** push frames.
- Keep **position-aware TTL freshness**: 90 seconds for addresses with positions **and for addresses where a snapshot has not yet arrived**; 15 minutes for confirmed empty (snapshot was, positions are 0) accounts. A uniform TTL of 90 seconds falsely "expires" empty accounts and initiates endless resubscriptions.
- Update the last message time **before** deduplication and processing, otherwise repeated snapshots are counted as silence and trigger unnecessary REST requests.
- The liveliness metric is — how many subscribed addresses sent something in the last 90 seconds / 15 minutes. Counting sent subscriptions proves nothing. Count the metric per each connection to see a dead shard.
---

## 4. Limits

### Table

| Limit | Value | Confidence / Source |
|---|---|---|
| Idle timeout server | 60 seconds without messages | high: documentation and live socket |
| WS-subscriptions per IP | around 1000 | medium (documentation), practically never reached the limit in real life |
| WS-connections per IP | records vary: 10 and 100; when exceeded, a `1008` close was observed. See ops-and-deploy.md §1.2 | low/medium, number not rechecked |
| Unique user addresses in user subscriptions | documented as **10**. Measured as **around 20 across all connections on an IP** (2026-07-17) | documentation / measurement |
| Limit per single connection | exists but **secondary**. The error message says "15", however, the error occurred even with 0 and 8 subscriptions | high |
| Ghost slot | around 60 seconds after unsubscribe on another connection or after reconnect | high, observed multiple times |
| Outgoing WS-messages per minute | limit exists; number not recorded here | blank |
| REST weight | WS-subscriptions **do not** consume REST weight (1200/minute per IP) | medium |
| Burst subscribe | large batches of subscribe seem to cause failures; working rule is no more than 5 resubscribes in a 30-second tick | low/medium |
### User-Tracking Limit Models: Which Ones Failed and Which One Applies

| Model | Result |
|---|---|
| "15 users on connection" (by error text) → sharding by 15 per socket | error exists, but model incorrect |
| Shard below 15 with extra slots for ghost | helps only partially |
| "About 30 total on IP" | rough estimate based on loaded IP, **incorrect** |
| Measurement 2026-07-17: **about 20 addresses per IP in total** (on a fresh IP out of 30 requested, about 20 are tracked, others get refusal). Pools from different IPs **are independent**: second egress-IP doubles the capacity, limit is not tied to server or account | **effective model** |

Practical rule: plan for documented **10 unique addresses per IP**. Anything above this consider an unguaranteed bonus (up to ~20), which HL may remove. More addresses — more egress-IPs or REST-polling.

Sharding by connections on one IP does not expand capacity. Model "15 per connection, thus sharding" is not applicable.
### Error Channel Error Lines

| Substring (compare in lowercase) | Value |
|---|---|
| `Cannot track more than 15 total users` | limit on user-tracking exceeded. Subscription looks accepted, but there will be **no** push from it |
| `already unsubscribed` | unsubscribe from a non-existent subscription |

Count `trackingRejects` **per each connection**. If rejections hit full connections while idle ones are silent, it's a connection limit. If rejection comes from an IP with few subscriptions in total but many overall, it's per-IP limit. You won't be able to distinguish this by the aggregate counter.

### How to Properly Measure Capacity

1. A snapshot on subscribe arrives **even for addresses that HL refused to track**. Measuring immediately after subscription inflates capacity.
2. Count only addresses that sent a **second** push, and wait several minutes. It's convenient to check with accounts having open positions, as they constantly push notifications.
3. Also monitor the number of `Cannot track…` by connections.
### How Does IP Overloading Develop

1. WS-slots or connections are filled; `1008` closures appear.
2. Some addresses stop receiving fresh snapshots: tracking failures and the number of stale addresses increase.
3. REST-fallback is activated, requests queue up in the throttler.
4. Delays grow, REST returns 429, WS closes by limit.

The symptom is invisible because REST-fallback "fixes" everything: data exists, just slowly and expensively.

---

## 5. Sharding, Priorities, and Slot Distribution
- **Slots are shared by all consumers on an IP.** When demand exceeds supply, HL silently rejects the excess, and slots go to **whoever subscribed first**, not to the most important consumers. A process that started earlier can occupy every slot, leaving none for critical subscriptions.
- **Explicitly set priorities with a hard cap on slots per consumer:**
  1. Subscriptions whose loss would cause a hard degradation (trading decisions depend on them);
  2. Accounts with open positions where pushes are needed for client stops (unrealizedPnl moves together with mark);
  3. All others on REST polling (soft degradation).
- **Do not use WS for what does not need pushes.** If protection is on the exchange (native SL/TP), there are no client stops, and decisions read from REST, then a WS slot for such an account only saves REST weight. Such accounts can be completely removed from WS and their slots allocated to more critical subscriptions.
- **Order of input list sets priorities:** important addresses first, overflow goes to REST. Log degradation (how many requested addresses are subscribed, who went to polling) **only when the set changes**. The same set on each tick is a no-op without traffic.
- **Do not subscribe inactive accounts,** they waste slots unnecessarily.
- **Egress bandwidths:** different consumers with different `localAddress` values get independent slot pools.
- **Two modules subscribed to one channel type at intersecting addresses should work through separate connections or via a common subscription manager with refcount.** Otherwise, deduplication by key `JSON.stringify(subscription)` results in one module's `unsubscribe` removing the other's subscription.
- **Stable set composition.** Sort addresses deterministically and deduplicate them to lowercase. If the same set "breathes" from tick to tick, it leads to extra unsubscribe/subscribe operations and ghost slots.
- **Keep demand below supply with a buffer for ghost slots.** When the set is close to the limit, `Cannot track` errors flow in, subscriptions silently die, reads flood into REST, and REST hits 429.
- A subscription set that depends on positions (position opened → push needed, closed → slot freed) should be reassembled **with debounce** (250 ms): a batch of own orders can result in dozens of FILLED statuses within milliseconds. Full synchronization with self-healing is only necessary when the configuration changes.
---

## 6. Ping/pong and Dead-Socket Detection

| Constant | Value | Purpose |
|---|---|---|
| `PING_MS` | 30 000 | server closes connection after 60 seconds of silence |
| `SILENCE_MS` | 65 000 (= ping×2 + 5 s) | no incoming frames → `ws.terminate()` |
| `handshakeTimeout` | 15 000 | timeout for hanging handshake |

- Start the ping timer after `open`, clear it on `close`.
- Filter out the `pong` frame, do not deliver it to consumers.
- Use `.unref?.()` for timers to prevent them from keeping the process alive during shutdown.
```ts
private startPing(): void {
  this.clearPing();
  this.pingTimer = setInterval(() => {
    if (this.lastMessageAt > 0 && Date.now() - this.lastMessageAt > SILENCE_MS) {
      try { this.ws?.terminate(); } catch { /* the close handler will reconnect */ }
      return;
    }
    if (this.ws?.readyState === WebSocket.OPEN) this.ws.send(JSON.stringify({ method: 'ping' }));
  }, PING_MS);
  this.pingTimer.unref?.();
}
```
---

## 7. Reconnect

### Constants

| Constant | Value |
|---|---|
| `RECONNECT_BASE_MS` | 1 000 |
| Multiplier | ×2 after each attempt |
| `RECONNECT_MAX_MS` | 30 000 |
| `RECONNECT_AFTER_LIMIT_MS` (close `1008`) | 60 000, not less than |
| Backoff reset | only after stable operation: socket `OPEN`, all subscriptions confirmed, no `serverError`. Options: 8 seconds or 30 seconds of stable work or first data frame (`channel !== 'subscriptionResponse'`) |
| Subscription confirmation timeout | 30 seconds without all ack → `subscription confirmation timeout` → `terminate()` |
### Rules

1. **Do not reset backoff to `open`.** HL sometimes accepts a socket and immediately closes it with code `1008`; resetting to `open` will cycle reconnects every 1 second. Another pitfall: an ack for one subscription may come together with `channel:error` for another. If you reset on ack, the partially subscribed socket will be reconnected every second indefinitely.
2. **Wait at least 60 seconds on `1008`.** Hanging connections (e.g., from previous dev-server restarts) hold a slot on the side of HL for about 60 seconds. Short retries only add to hanging connections.
3. **Resubscribe after reconnecting.** Set the state to `connected` only when all acks have arrived (the set of pending acks is empty).
4. **Acknowledge only by an exact match of `channel:'subscriptionResponse'` with `data.method === 'subscribe'` and matching `subscription`.** Acks for unsubscribes or stray data frames should not make a partially subscribed socket healthy. In a multiplexed socket, a frame from one user should not confirm a subscription for another.
5. **The key for ack is** `JSON.stringify({ type, coin, user: user?.toLowerCase() || undefined })`, because the server may return an address in a different case.
6. **Start every handler with `if (this.ws !== ws) return;`**. Otherwise, a close from an old, already replaced socket (after `terminate`) will reset the state of the new one and trigger a double reconnect.
7. **HL holds slots for about 60 seconds after a reconnect for old connections** (ghost). Resubscribe with backoff and do not hammer immediately.
8. **The history frame on reconnect** (`userFills` with `isSnapshot`) counts, but positions and orders should be verified via REST asynchronously.

### Channel `error` and fate of the socket
- **Log `channel:error` always**, this is an obligatory rule. Without it, the subscription failure will be completely invisible: the subscription "succeeds," there are no data, and everything quietly goes into REST.
- HL **may leave TCP open** after a failure. If the socket is not closed, `isConnected` will lie infinitely, and the rejected set will never resubscribe. A simple solution is to reset the confirmation flag and do `ws.terminate()` on any `error`, rather than graceful close: graceful close waits for a close timeout `ws` from the same unhealthy peer.
- **If the error is due to capacity (`Cannot track…`), reconnection won't help:** per-IP pool, but reconnection itself generates ghost slots for 60 seconds. Leave the socket, count the failure, reduce the set or migrate excess to REST. This conclusion comes from a combination of facts above and has not been verified separately.

### Minimal Client (raw `ws`)

```ts
import WebSocket from 'ws';

const WS_URL = 'wss://api.hyperliquid.xyz/ws';
const PING_MS = 30_000;
const SILENCE_MS = 65_000;
const RECONNECT_BASE_MS = 1_000;
const RECONNECT_MAX_MS = 30_000;
const RECONNECT_AFTER_LIMIT_MS = 60_000;
const STABLE_MS = 8_000;

type Sub =
  | { type: 'l2Book'; coin: string }
  | { type: 'userFills'; user: string; aggregateByTime?: boolean }
  | { type: 'orderUpdates'; user: string }
  | { type: 'allDexsClearinghouseState'; user: string };

const ackKey = (s: { type: string; coin?: string; user?: string }) =>
  JSON.stringify({ type: s.type, coin: s.coin, user: s.user?.toLowerCase() || undefined });

export class HlWs {
  private ws: WebSocket | null = null;
  private subs = new Map<string, Sub>();
  private acks = new Set<string>();
  private reconnectDelay = RECONNECT_BASE_MS;
  private lastMessageAt = 0;
  private shouldRun = false;
  private pingTimer?: NodeJS.Timeout;
  private stableTimer?: NodeJS.Timeout;
  private handlers: Array<(channel: string, data: unknown) => void> = [];
  public trackingRejects = 0;

  constructor(private localAddress?: string) {}

  onMessage(h: (channel: string, data: unknown) => void) { this.handlers.push(h); }
  get isHealthy() {
    return this.ws?.readyState === WebSocket.OPEN && [...this.subs.keys()].every((k) => this.acks.has(k));
  }

  start() { this.shouldRun = true; this.connect(); }
  stop() { this.shouldRun = false; this.clearTimers(); this.ws?.terminate(); this.ws = null; }

  subscribe(sub: Sub) {
    const k = ackKey(sub); if (this.subs.has(k)) return;
    this.subs.set(k, sub);
    this.send({ method: 'subscribe', subscription: sub });
  }
  unsubscribe(sub: Sub) {
    const k = ackKey(sub); if (!this.subs.delete(k)) return;
    this.acks.delete(k);
    this.send({ method: 'unsubscribe', subscription: sub }); // on THIS SAME socket the slot is freed immediately
  }

  private send(obj: unknown) {
    if (this.ws?.readyState === WebSocket.OPEN) this.ws.send(JSON.stringify(obj));
  }

  private connect() {
    const ws = new WebSocket(WS_URL, { perMessageDeflate: false, handshakeTimeout: 15_000, localAddress: this.localAddress });
    this.ws = ws;

    ws.on('open', () => {
      if (this.ws !== ws) return;
      this.lastMessageAt = Date.now();
      for (const sub of this.subs.values()) ws.send(JSON.stringify({ method: 'subscribe', subscription: sub }));
      this.startPing();
      this.stableTimer = setTimeout(() => {           // reset backoff only on a stable socket
        if (this.ws === ws && this.isHealthy) this.reconnectDelay = RECONNECT_BASE_MS;
      }, STABLE_MS);
      this.stableTimer.unref?.();
    });

    ws.on('message', (raw) => {
      if (this.ws !== ws) return;
      this.lastMessageAt = Date.now();
      let msg: { channel: string; data?: any };
      try { msg = JSON.parse(raw.toString()); } catch { console.warn('[ws] bad json'); return; } // don't break the socket
      if (msg.channel === 'pong') return;
      if (msg.channel === 'subscriptionResponse') {
        if (msg.data?.method === 'subscribe' && msg.data.subscription) this.acks.add(ackKey(msg.data.subscription));
        return;
      }
      if (msg.channel === 'error') {
        const text = String(msg.data ?? '');
        console.error('[ws] server error:', text);     // MUST BE LOGGED
        if (text.toLowerCase().includes('cannot track more than')) { this.trackingRejects++; return; }
        try { ws.terminate(); } catch {}                 // other errors: the socket may 'hang' open
        return;
      }
      for (const h of this.handlers) {
        try { h(msg.channel, msg.data); } catch (e) { console.error('[ws] handler threw', e); }
      }
    });

    ws.on('error', () => { /* always followed by close */ });

    ws.on('close', (code: number) => {
      if (this.ws !== ws) return;
      this.clearTimers(); this.ws = null; this.acks.clear();
      if (!this.shouldRun) return;
      if (code === 1008) this.reconnectDelay = Math.max(this.reconnectDelay, RECONNECT_AFTER_LIMIT_MS);
      const delay = this.reconnectDelay;
      this.reconnectDelay = Math.min(this.reconnectDelay * 2, RECONNECT_MAX_MS);
      setTimeout(() => this.connect(), delay).unref?.();
    });
  }

  private startPing() {
    this.pingTimer = setInterval(() => {
      if (this.lastMessageAt > 0 && Date.now() - this.lastMessageAt > SILENCE_MS) {
        try { this.ws?.terminate(); } catch {}
        return;
      }
      this.send({ method: 'ping' });
    }, PING_MS);
    this.pingTimer.unref?.();
  }
  private clearTimers() { clearInterval(this.pingTimer); clearTimeout(this.stableTimer); }
}
```
If WS is enabled but not all subscriptions are confirmed, this is a degradation to REST rather than a failure. **Do not log raw WS errors and addresses to public logs or metrics**, only states and counters. In the URL for logs, remove username/password/query.

### SDK `@nktkas/hyperliquid`

In 0.27.1, by default, `WebSocketTransport` has `maxRetries=3`. After exhausting attempts, the socket is closed **forever** and will not reconnect. For a bot that runs for days, this is not suitable: you need your own minimal `ws` client (as above) with infinite reconnection.

---

## 8. Self-healing: Silent Address Resubscription

Sending one subscribe request is insufficient: subscriptions can silently disappear, especially after a reconnection. A cycle is needed that checks the freshness of each address and resubscribes to stale ones.
| Constant | Value |
|---|---|
| Reconcile-tick | 30 s (and additionally when configuration or client changes) |
| `WS_SILENT_STALE_MS` / `SNAPSHOT_STALE_MS` | 90 000: addresses with positions or still without snapshot |
| `EMPTY_SNAPSHOT_STALE_MS` | 15 min: confirmed empty |
| `RESUB_AFTER_STALE_TICKS` | 2 |
| `MAX_RESUB_BACKOFF_TICKS` | 20 (≈10 min) |
| `MAX_RESUBS_PER_TICK` | 5 |
| Required ticks before resubscription | `min(2 * 2^attempts, 20)` |
| Log statistics | once every 60 s (or per tick) |

Rules:
- **Resubscribe on THE SAME CONNECTION:** `unsubscribe` + `subscribe`. Switching to another shard is forbidden, see pitfalls of ghost-storm.
- **Backoff is mandatory.** Without it, "silent" addresses will resubscribe infinitely, and resubscriptions create ghost-slots and kick live subscriptions out: only a part of subscriptions remains fresh, and `Cannot track` errors start appearing.
- **Self-healing should be run from the timer,** not from the event path. The first snapshot comes in 0.5–2 s, while calls from the event path come in dozens per millisecond: the address will "expire" before the first snapshot arrives, get resubscribed, occupy a ghost-slot, and a storm starts. Backoff counters are designed for 30-second ticks.
- A fresh message resets both counters (`staleTicks`, `resubAttempts`).
- Side effect for silent accounts: resubscription works as cheap WS-polling because the snapshot comes with every subscribe.
- When an address is removed from the set: `unsubscribe` on its shard and clear **all** structures related to the address (shard map, subscribed, staleTicks, resubAttempts, lastSeen, hasPositions).
```ts
// once every 30 seconds
let resubs = 0;
for (const wallet of this.subscribed) {
  const conn = this.connByWallet.get(wallet)!;
  if (!conn.isHealthy) continue;
  const lastSeen = this.lastSeenByWallet.get(wallet) ?? 0;
  // hasPositions === undefined → snapshot hasn't been received yet → short threshold
  const staleThreshold = this.hasPositionsByWallet.get(wallet) === false ? EMPTY_SNAPSHOT_STALE_MS : WS_SILENT_STALE_MS;
  if (now - lastSeen <= staleThreshold) { this.staleTicks.delete(wallet); this.resubAttempts.delete(wallet); continue; }
  const ticks = (this.staleTicks.get(wallet) ?? 0) + 1; this.staleTicks.set(wallet, ticks);
  const attempts = this.resubAttempts.get(wallet) ?? 0;
  const requiredTicks = Math.min(RESUB_AFTER_STALE_TICKS * 2 ** attempts, MAX_RESUB_BACKOFF_TICKS);
  if (ticks >= requiredTicks && resubs < MAX_RESUBS_PER_TICK) {
    resubs++; this.staleTicks.set(wallet, 0); this.resubAttempts.set(wallet, attempts + 1);
    conn.unsubscribe({ type: 'allDexsClearinghouseState', user: wallet }); // same socket!
    conn.subscribe({ type: 'allDexsClearinghouseState', user: wallet });
  }
}
```
`hasPositionsByWallet` fill only after successfully parsed snapshot: `positions.length > 0` by main or xyz.

### Health-String

Log the number of subscriptions, number of addresses with fresh data, number of snapshots with positions and among them with filled margin, number of live connections once per tick.
- Healthy state: all subscribed addresses have fresh data, all connections are alive, payload contains margin for all snapshots with positions.
- If there are noticeably more subscriptions than addresses with fresh data, this is a clear sign of silently rejected or lost subscriptions: a plateau where only part of the subscriptions remain fresh can be seen exactly like that.
- If snapshots with margin are consistently fewer than snapshots with positions, for some accounts payload comes without `totalMarginUsed`. Solutions for them should go to REST; sort out the form of payload before trusting WS.

---

## 9. Parsing Messages
- Invalid JSON log and skip. **Do not close the connection**: due to a single bad payload, otherwise there will be a reconnection and loss of snapshots for all subscriptions.
- Wrap each handler in `try/catch` separately, so that a broken handler does not interfere with others.
- Log initial subscription confirmations with preview (e.g., first 10 characters out of 200). Unknown channels log with message number.
- **Meta in the hot path.** If the parser calls `getMeta()` on every message (resolving `asset → coin`, `szDecimals`), this should be read from process cache with single-flight. Otherwise, each cache miss at the edge of TTL turns the WS stream into a burst of heavy REST requests ("meta stampede").
- **Trust verification before financial decision:**
  - Positions exist but `totalMarginUsed == 0` → snapshot is unreliable, rollback to REST. Otherwise `marginRatio = 0`, and margin risk check will not work;
  - `wsBalance !== null && (positions === null || positions.length === 0 || wsBalance > 0)`: open positions with zero `accountValue` mean that the snapshot is stale or incomplete → REST;
  - Snapshot older than 90s → REST.
- **Conservative defaults.** If REST for spot stablecoins fails, assume 0: account value lower, margin ratio higher, and risk check will more likely deny action.
- **`leverage.value` sometimes transiently equals 0 or is absent.** Do not calculate values dependent on leverage (ROE etc.) on this tick; leverage will return with the next snapshot.
- **Position with `szi=0`** (encountered in HIP-3 xyz data): snapshots may contain entries with zero size. Filter them out (as in parser §2), do not consider them as open positions. Code that compares snapshots otherwise will see a phantom opening of zero size, and the real opening later will not be recognized.
- Address `user` from payload convert to lowercase and use as key for all caches.
---

## 10. Delays and Races

| Phenomenon | Value | What to Do |
|---|---|---|
| First snapshot after subscribe | 0.5–2 s | do not consider the address dead until the first reconcile tick |
| `orderUpdates` `open` against HTTP response | WS arrives **before** | "adopt" the order from WS and merge with the response |
| Position closure in WS-snapshot | window in seconds, while a closed position is still visible | account for your recently sent orders |
| Single position change in WS-snapshot and REST-response | discrepancy ~1–3 s | make the handling of changes idempotent |
| Delayed reaction only on REST polling | up to poll interval + tick time | use WS `userFills` as a trigger → residual lag ≈ debounce 300 ms + tick time |
| Fan of fills | tens of FILLED in tens of milliseconds | trailing debounce 250–300 ms |
### Pattern «WS as Alarm Clock»

```ts
let debounce: NodeJS.Timeout | undefined;
ws.onMessage((channel) => {
  if (channel !== 'userFills') return;
  clearTimeout(debounce);
  debounce = setTimeout(() => scheduleSync(), 300); // collapses burst partial fills
});

async function sleepUntilNextTick(ms: number) {
  // wait MIN(WS-trigger, normal interval); trigger flag is reset BEFORE tick
}
```
WS here is only a **trigger**: the tick still reads fresh state via REST. The default polling interval remains as a fallback in case of WS disconnection.

Limitation of the pattern: the alarm after 300 ms from **the first** aggressive limit order fill triggers exactly within the window, while the order's remaining amount is still visible in the book. Solutions that depend on the set of open orders cannot be made based on a single snapshot immediately after the alarm: the snapshot will show a transient order.

---

## 11. Architectural Patterns

1. **WS — speed, REST — truth.** Position from REST `clearinghouseState` (e.g., every 2 seconds), between snapshots it is moved by WS fills. Orders: local book from responses to placing and canceling orders plus WS events, reconciliation through REST `openOrders` once every 10 seconds, **after reconciliation the exchange rights are verified**.
2. **Unhealthy WS on its own events — reason for pause.** Without `userFills`/`orderUpdates`, the local book and position "go blind". If the socket with these subscriptions is unhealthy longer than a threshold (e.g., 5 seconds), put order placement on hold. After recovery, immediately do reconciliation of `openOrders` and an extra snapshot of `clearinghouseState`. Position without REST confirmation for longer than a threshold (e.g., 15 seconds) — also pause.
3. **WS-snapshot as cache read.** Getter returns `null` if there is no snapshot (only subscriptions, reconnect, disconnection) or it has expired, and the caller falls back to throttled REST. Keep the WS-path switch in configuration to return to REST without changing code.
4. **Hybrid WS + rare REST-polling.** Parallel REST-polling as insurance passes accounts with fresh WS data; change handling is idempotent (§10).
5. **Cold start and scripts.** Standalone script that imports the service without starting WS gets an empty WS-cache. Any path on the WS-cache (especially closing positions) should be able to take state through REST.
6. **Degraded — working state.** Upon disconnection of WS, the client switches to REST; health shows `degraded`; this is degradation, not a failure.
7. **Stub instead of address in config** (`0x...` from example env) in subscription `userFills`/`orderUpdates` receives `channel:error`, and starts an infinite reconnect loop. Validate addresses and keys with regex **always**, when they are defined, even in dry-run. For dry-run without counting, leave the variable empty.
---

## 12. Pitfalls

| # | What Breaks | Why | How to Do It Right |
|---|---|---|---|
| 1 | Subscription "accepted", no data, everything quietly goes into REST: readings stall, REST weight accumulates | user-tracking limit; failure only comes in `channel:error` | log errors, count connection failures, keep demand below supply |
| 2 | Sharding by connections does not add capacity | per-IP limit (~20), not 15 per connection | more capacity — only new egress-IP or REST |
| 3 | Self-sustaining failure storm at demand below the ceiling: only a part of addresses stays fresh | resubscribing to another connection leaves a ghost slot for ~60s → failure from neighbor → stale → new resubscription | re-subscribe only on the same socket; backoff; no more than 5 per tick |
| 4 | Subscriptions are restored only partially after reconnect (observed after close reason `Expired`) | lingering connection holds slots for ~60s | self-healing with backoff, wait for slot release |
| 5 | Important subscriptions remain without slots | slots go to the first subscriber by IP | strict limits and priorities |
| 6 | Shards are right at the limit: stream of `Cannot track` and 429 from REST | no buffer for ghost slots; stale snapshots flood into REST | slot buffer, prioritization, fewer WS consumers |
| 7 | Fresh snapshot plateau is below the number of subscriptions and infinite resubscription | single TTL of 90s, but empty accounts do not push | position-aware TTL: 90s / 15min |
| 8 | Capacity measurement is overestimated | subscription snapshot goes to abandoned addresses | count second push, wait a minute |
| 9 | Decision made on already closed position | WS snapshot lags by seconds after closing | account for own in-flight orders, do not trust snapshot immediately after your order |
| 10 | One position change processed twice | WS and REST see it with 1-3s delay | idempotent processing |
| 11 | Burst of heavy metadata REST requests | `getMeta()` per each WS frame without single-flight | process cache + single-flight |
| 12 | Infinite reconnect every second | backoff resets on `open` or ack, HL closes with `1008` or sends `error` to part of the set | reset only after stable work (8-30s) with all acks |
| 13 | Avalanche of hanging connections | short retries on `1008` | wait at least 60s |
| 14 | `isConnected` "lies" forever | TCP stays open after `error` from HL | for non-capacity errors, do `terminate()` |
| 15 | Double reconnect, state reset of new socket | `close` old socket after `terminate` | `if (this.ws !== ws) return;` in each handler |
| 16 | Client is "unhealthy" forever | ack compared with address in different case | key ack with lowercase `user` |
| 17 | Bot stops receiving data after some time | WS transport dies completely after 3 failed reconnects for SDK | custom client with infinite reconnect |
| 18 | False alert "disconnection" on a quiet market | `l2Book` pushes only on changes | detect all frames every 65s + REST orderbook |
| 19 | Position or PnL doubled after reconnect | `userFills` frame `isSnapshot:true` applied as new fills | snapshot only in accounting (`time >= sessionStart`, dedup `tid`), position from REST |
| 20 | One module unsubscribes data of another | common client, dedup by `JSON.stringify(subscription)` | separate connections or refcount |
| 21 | Excessive churn and ghost slots | order entry set changes between ticks | deterministic sort + dedup |
| 22 | Ghost storm from reassembly of set on every FILLED | dozens of calls per millisecond trigger self-healing | debounce 250ms, self-healing only from 30-second timer |
| 23 | Unknown oid in local book («snatched from WS, it was not in the local book») | `orderUpdates` precedes HTTP response | adopt + merge with response |
| 24 | Counter of post-only failures ×2 | failure comes both in the response and in `orderUpdates` `badAloPxRejected` | dedup by oid or event |
| 25 | Money resolution from snapshot without margin → risk check based on margin did not trigger | payload without `totalMarginUsed` for open positions | sanity-check → `null` → REST |
| 26 | Zero position taken as open (HIP-3 xyz) | record with `szi=0` in the snapshot | filter out `szi=0` |
| 27 | Incorrect ROE or other leverage value | transient `leverage=0` | skip tick, leverage will return on next snapshot |
| 28 | Spot balance «vanished» from equity | no WS spot available | REST `spotClearinghouseState` with TTL |
| 29 | Resolution for set of open orders taken from snapshot immediately after the alarm | trigger within 300 ms window while order remains in book | confirm with multiple snapshots (not verified, see open questions) |
| 30 | Infinite reconnection during dry-run | placeholder address from example env → `error` | validate with regex always |

---

## 13. Open Questions / Not Verified
- **Exact Limits.** Documented 10 unique user-addresses, text error says "15", measured around 20 per IP (2026-07-17). Earlier estimates of "15 per connection" and "around 30 per IP" were contradicted by the same-day measurement. It is unclear what exactly the number 15 in the text error counts and how stable the output is above 10. HL may tighten the limit to the documented one.
- **Number of WS-connections per IP.** Records differ: 10 against 100 (ops-and-deploy.md §1.2); official rate limits page not rechecked. There is no exact limit on outgoing WS-messages per minute here.
- **~1000 subscriptions per IP** — average confidence, never reached this ceiling in live use.
- **`spotState` in `allDexsClearinghouseState`.** The type `spotState: { totalRawUsd }` appears in documentation clients, but with a note that its shape is checked at runtime. A later check (2026-06-11) found that spot balances are not available over WS and must be read through REST. This conflict is resolved in favor of REST; the presence and meaning of `spotState` are not verified.
- **Frequency of push `allDexsClearinghouseState`.** "On every mark movement" (average confidence) or "every few seconds" — exact interval not measured.
- **Statuses of `orderUpdates`** are the same as for `orderStatus`/`historicalOrders` (orders.md §12.2): `open`, `filled`, `canceled`, `triggered`, `rejected`, `marginCanceled`, `vaultWithdrawalCanceled`, `openInterestCapCanceled`, `selfTradeCanceled`, `reduceOnlyCanceled`, `siblingFilledCanceled`, `delistedCanceled`, `liquidatedCanceled`, `scheduledCancel` and any `…Rejected`. Confirmed live on socket for `open` and `badAloPxRejected`; others by coincidence with snapshots of `historicalOrders`.
- **Not verified live:** `trades`, `allMids`, `candle`, `bbo` (the shapes in §2 come from documentation and were not verified against a live socket), `userEvents`, `userFundings`, the `webData2` shape, and sending actions (`post`) through WS. Verified on a live socket on 2026-09-23: the set of `activeAssetCtx.ctx` keys (table §2), and `post` responses to invalid info requests — `500 Internal Server Error`, `data: null`, and an `error` frame echoing the envelope (sdk-and-api.md §3.6). Recording one-second candles requires `trades` with deduplication by `tid`; reconnect behavior (whether a recent-trades snapshot exists) is not verified.
- **SDK 0.27.x.** No verified snippet for WS-subscriptions through `@nktkas/hyperliquid`. Can `maxRetries` be set to infinity — not verified.
- **Backoff reset:** the options are "first data frame," "8 s of stability," and "30 s of stability"; which is better has not been compared.
- **Reaction on `channel:error`.** `terminate()` works for any error as long as the address count stays within 10. For capacity failures, reconnection, judging by ghost-slot mechanics, may be harmful. Not separately verified.
- **Burst subscribe.** What HL drops subscription batches — phrased "seems like", speed limit of subscribe not known.
- **Close reason `Expired`** — semantics not clarified.
- **Solutions for the set of open orders immediately after WS-alarm** (transient window 300 ms): recipe "confirm with multiple snapshots" not verified.
---

Knowledge slice — 2026-09, dates of individual verifications — in the text. API HL changes, so recheck limits and response forms.

---

<!-- license-footer -->
_© markpaper authors. Licensed under [CC BY 4.0](LICENSE.md): when publishing or adapting, credit “markpaper — Hyperliquid knowledge base” and provide links to the original and the license._
All files