ADR-003: BotCore as the state layer, peer to ArbitrageEngine¶
Status: accepted. Recorded during the BotCore/ArbitrageEngine separation-of-concerns grilling, June 2026. Implemented in full by Plan 100 (Slices 1–5): V2/V3/V4 state consolidated into BotCore; the V2BlockEngine/V3BlockEngine/V4BlockEngine are dissolved (single live pool-state owner; engine-then-core lock order); the legacy RustPoolCache/ArbPoolCacheAdapter mirror is deleted (Slice 4 Option D); PyToken is completed (Slice 5). Supersedes the implicit arrangement where each block engine owned a private pool-state HashMap and BotCore sat unused.
Context¶
BotCore (rust/src/bot_core/mod.rs) was designed as the single Rust owner of all runtime pool/token state, with thin PyPool/PyToken handles over Arc<Mutex<BotCore>>. It is the only Rust struct with reorg-rollback journals (v2/v3_restore_before_block, scalar + per-tick priors) and the only one that ever returned PyPool/PyToken handles.
ArbitrageEngine (rust/src/optimizers/arb_engine/) was prototyped as a mixed V2/V3/V4 arbitrage engine — a tracer bullet. In production it grew to own pool state directly: each of V2BlockEngine, V3BlockEngine, V4BlockEngine holds a private HashMap of pool state (reserves/tick_data/sqrt_price), and the PyO3 wrapper exposes ~40 methods, most of which are state registration/mutation/snapshot/verification rather than solving. BotCore is instantiated nowhere (production uses ArbitrageEngine; the library solve path uses RustPoolCache via ArbPoolCacheAdapter, a parallel Rust backend). Its V3 calculate_tokens_out is a stub returning U256::ZERO.
The result is two parallel Rust state implementations of the same V2/V3 pools, plus a third live copy inside RustPoolCache for the legacy solver path.
Decision¶
Adopt Option 1 — BotCore as a peer module. BotCore becomes the single owner of pool and token state (V2, V3, and a new V4 variant). ArbitrageEngine keeps path registry, solver dispatch, result batching, the pump, and diagnostics — and reads/mutates state through BotCore via a shared Arc<Mutex<BotCore>>, not through private per-engine stores.
The two lock in a fixed order when nested: engine-then-core. The pump (single Tokio task) holds the engine lock for its per-block coordination; resolve-time briefly acquires the core lock inside the engine lock to re-derive solve-ready hop states. No code path acquires them in the opposite order.
PyPool and PyToken become the construction entry point and per-pool read handle — independent of the engine — mirroring ADR-001’s “thin handles over Rust-owned state.”
Considered options¶
Option 2 — State as a field of
ArbitrageEngine. One lock (Mutex<ArbitrageEngine>) covers state + solve coordination. Rejected: makes BotCore unusable without the engine, so the legacyArbPoolCacheAdapter+RustPoolCachepath can only retire by deletion of the Python side, leaving a parallel Rust store. Entangles state with pump coordination.Option 3 —
BotCoreas a field of the engine, same lock. Same single-lock benefit as Option 2 with a cleaner internal seam. Still rejected for the same reason as Option 2: the legacy path cannot migrate onto it.Option 1 — peer modules. Costs lock-ordering discipline and an extra critical section at resolve time. Chosen because (a) it is the only topology that retires the legacy
RustPoolCachesecond Rust backend at all (see “Legacy solver path retirement” below — the retirement is by deletion, not migration), (b) it makes a future non-arb consumer of Rust state (sync calc, Curve port) possible without standing up the whole pump, and (c) the deadlock surface is empty in practice: the pump is single-writer on the engine lock during the hot loop; the only Python contender islatest_results, which reads engine-localself.resultsand does not touch the core lock.
Core-lock placement: per-call (Option A)¶
Within the engine-then-core lock order, the core (BotCore) lock is acquired per pump step, inline, with no engine-side mutation buffer:
Per WS log (
apply_log): pump holds engine lock, briefly nests the core lock, mutates BotCore directly (no copy, no buffering), releases both. Inserts affected pool IDs into the engine’s dirty sets exactly as today. BotCore is current the instantapply_logreturns — the literal eager-processing invariant.Per coalesced solve (
solve_dirty): pump holds engine lock, takes the core lock once for the re-derive of all N affected paths (single consistent snapshot, no re-acquire thrash), releases, thensolve_pathruns pure&selfover the per-path resolved cache exactly as today.
This is Option A over Option B (engine-side transient mutation buffer drained only by solve_dirty). A was chosen because:
The performance argument for B was illusory at mainnet scale — the uncontended
parking_lot::Mutex(~25ns) × ~15 relevant logs/block ≈ 375ns/block against a 12s block time.solve_dirtyunder A takes the core lock once for the whole re-derive and sees a consistent snapshot, exactly as B would — the re-derive benefit is not B-exclusive.B reintroduces a transient pool-state copy on the engine, however short-lived, which muddies ADR-003’s clean “the engine owns no pool state” claim and adds drain invariants. A preserves it literally.
Lock-ordering rule¶
Engine-then-core is the only direction that ever nests, and the two locks are taken by disjoint caller sets.
Python-facing
PyBotCoremethods (register_v2_pool,update_v2_pool,get_pool,calculate_tokens_out, etc.) take the core lock alone — they never call into the engine.The pump and engine methods take engine-then-core when they need both.
No code path holds the core lock and then calls into the engine. Such a path would invert the order and deadlock against an engine-then-core caller. A future Python method that wants both must take the engine lock first.
The deadlock surface is empty in practice: the pump is single-writer on the engine lock during the hot loop; the only Python contender, latest_results, reads engine-local self.results and does not touch the core lock.
Reorg detection and restore (Option α)¶
Detection lives on the pump (it owns WS event ordering already). The pump reads the canonical reorg signal — the removed: bool field Alloy exposes on every eth_subscribe log — rather than inferring forks from block-number comparison. removed: true means the log was orphaned; removed: false means it’s on the canonical fork (either fresh or re-emitted after a reorg). Detection by removed-flag is authoritative: it has no false positives from out-of-order delivery and catches the case where a removed log’s block number is still ≥ last_solved_block (which a block-number heuristic would silently drop).
Restore coordination lives on the engine (it spans the solver’s derived state and BotCore’s pools). On detecting a reorg, the pump calls engine.handle_reorg(target_block) under the engine lock, which:
Acquires the core lock (engine-then-core ordering) and calls
core.restore_all_pools_before_block(target)— per-poolReorgJournal::restore_before_block(target). Pools untouched by the reorg have no delta at/aftertarget, so this is a no-op per pool; touched pools get scalar + per-tick priors restored. Idempotent.Invalidates
path_resolvedfor all paths (marks every path dirty) — derived state was built from pre-restore pool state.The next
solve_dirtyre-derives and re-solves naturally;compute_diff_and_sendemitsexpired/updated/removeddiffs against thedeliveredset. Python sees forked paths vanish from batches without an “un-receive” —deliveredstays the truth of what Python has seen.
Journal depth is user-configurable, default 1 mainnet epoch (32 blocks). A reorg deeper than the journal cannot be restored from deltas — that path is fail-stop the pump with a diagnostic (a >32-block reorg on mainnet is a Severity-1 event where auto-recovery can mask bigger problems). The user-configurable depth lets operators on other chains (or with different risk tolerance) raise or lower the bound.
Mid-block reorg window is accepted as inherent to eager processing. If apply_log applied logs from block N before the next WS message reveals N was reorged, the journal’s delta-pushing apply_* (ADR-003) makes those applies recoverable — restore_before_block undoes them — but the window exists. Deferring apply_log until block finality would close the window but destroys the zero-latency property that is the whole point of the eager architecture. Accepted as a known bound; recovery via the journal is first-class rather than “restart from snapshot.”
Legacy solver path retirement: delete, not migrate¶
The legacy library solve path is ArbitragePath (Python) → ArbSolver (Python, holds a Rust RustPoolCache) → RustPoolCache (Rust PyO3 surface that mirrors Python pool reserves and clones per-path hop state). ArbPoolCacheAdapter (Python) subscribes to Python pool state updates and pushes reserves+fee into the Rust RustPoolCache. This is a second Rust backend alongside the production BotCore/ArbitrageEngine path.
Decision: delete the mirror. Specifically:
Rust:
RustPoolCache,RustIntHopState,RustArbResultPyO3 classes deleted frommobius_py.rs.Python:
ArbPoolCacheAdapterdeleted.ArbSolver’s registered-path surface (register_pool,update_pool,remove_pool,register_path,update_path,update_all_paths,remove_path,solve_registered,solve_registered_ints,solve_cached,solve_cached_batch,get_pool_cache) deleted.ArbSolver.solve(SolveInput)retained on the pure-Python f64 path (the Rust fast paths were already removed).Python:
ArbitragePathunchanged — it buildsSolveInputfrom pool state at solve time and callssolver.solve(...), which never depended on the mirror.Rust:
PyBotCore/PyPool/PyTokengain nothing from this retirement — they serve the production path, not the legacy one.
Considered options¶
M — migrate the mirror onto
BotCore.ArbPoolCacheAdapterwould retarget fromRustPoolCachetoPyBotCore; the registered-path solve APIs would land onBotCore. Rejected: this re-pollutesBotCorewith solver-shaped concerns (path registry, registered-path solve) that ADR-003 keeps onArbitrageEngine—BotCoreowns state, the engine owns solving. Migration isn’t a clean retirement, it’s moving the pollution to a new home. Worse, it would leave two Rust backends reading the same conceptual pool state through two different Python adapters.K — keep as-is. Leaves two parallel Rust backends indefinitely; the future architecture (one Rust core, one pump-driven engine, thin Python handles) never arrives. Rejected.
D — delete the mirror. Chosen. After deletion, Rust state lives in exactly one place (
BotCore+ itsLiquidityMaps); Python reads through one handle family (PyPool/PyToken/PyBotCore); Python solves through one engine (ArbitrageEngine) or through the pure-PythonArbSolver.solve(SolveInput)for library consumers without a mirror.
Why D serves the long-term goal (Rust core + thin Python interface)¶
M and K both preserve the second Rust backend indefinitely. D removes it today. When Curve ports to Rust, the question “does new Curve state live in BotCore’s LiquidityMap or in RustPoolCache?” has only one answer under D; under M/K it’s ambiguous and the easier path is to keep using the mirror — perpetuating the split. D makes the architecture converge under its own gravity: one Rust core, one engine, thin handles, and a pure-Python fallback left in place for not-yet-ported families until their Rust-native equivalents under ArbitrageEngine arrive (at which point ArbSolver.solve(SolveInput) itself deletes).
Cost acknowledged¶
ArbitragePath consumers lose the unused Rust-accelerated registered-path solve path. The drop is from a speed nobody gets (no library caller exercises solve_registered/solve_cached — verified: ArbitragePath calls only solver.solve(SolveInput), the pure-Python f64 path) to a speed they were always going to get once the mirror retired. Production (examples/eth_backrun_v2_v3_v4_rust.py) uses ArbitrageEngine directly and never touches RustPoolCache/ArbPoolCacheAdapter/ArbSolver.
Deletion-order caveat for the plan¶
ArbPoolCacheAdapter’s isinstance(pool, CacheablePool) check and the CacheablePool protocol (reserves_for_cache/fee_for_cache) are used only by the adapter (verified). The V2/V3/Aerodrome pool classes’ reserves_for_cache/fee_for_cache methods may be exposed for other reasons — the plan’s deletion-test step must verify before removing the protocol itself.
Token state: Rust owns metadata, Python owns the price oracle¶
Token state splits along the same line as pool state under ADR-003: Rust owns the state it computes on; Python owns what it orchestrates.
Rust-owned (
BotCore.tokens: HashMap<Address, TokenEntry>): address, decimals, symbol, name — the immutable metadata that Rust-core computations need without a Python round-trip. Two such computations drive this: (1) decimal normalization for cross-token profit comparison and multi-hop output reconciliation, (2) token-equivalence rules in path validation (e.g. WETH == EtherPlaceholder on chains with native ETH).PyTokenis the thin handle reading this state — the analog ofPyPoolfor tokens, not a stub.Python-owned (
Erc20Tokenwrapping aPyToken): the price oracle (ChainlinkPriceContract), an I/O construct with its own subscription/refresh lifecycle that cannot move to Rust.Erc20Tokenbecomes an orchestration layer over aPyTokenhandle, adding price-oracle + display concerns — the same split asPyPool(Rust: state+math) vs.PyBotCore/Bot(Python: I/O orchestration).
build_erc20token stays Python-side as the construction entry point — it fetches metadata from RPC, constructs the PyToken in BotCore, and wraps it with the price oracle.
Considered options¶
T2 — delete
BotCore.tokens/PyTokenentirely. Considered and rejected during grilling. Rust-core computation wants decimal normalization and token-equivalence rules; without Rust-owned token metadata those computations either round-trip Python on every comparison or simply can’t exist in Rust. Pulls against the future architecture (the whole point of Rust core is avoiding per-computation Python round-trips).T1 — keep as-is structurally, leave
PyTokenasaddress-only stub. Rejected:addressis already onPoolEntry(token0/token1), so a stub returning onlyaddressadds nothing. The interface has to be completed (decimals/symbol/namegetters) or the struct is dead weight.T3 — complete
PyTokenas a real read handle (chosen). Rust owns token metadata as state;PyTokenreads it; Python’sErc20Tokenwraps the handle with the price-oracle I/O layer. Same line as the pool-side split.
Consequences¶
The block engines lose their private pool-state
HashMaps. They are dissolved: V2’s engine (no non-state concerns) deletes entirely; the V3/V4 buffer-apply-verify trinity moves intoBotCoreas a first-classLiquidityMapconcept (one per CL family, keyed byAddress/(Address, PoolId)). Accurate pool state is a cross-cutting concern, not a solver-engine concern — diagnostics, verification, classifiers, and a future Curve port all consume it without going through the solve engine.ArbitrageEngine’sapply_logbecomes a consumer ofBotCore’sLiquidityMap, not an owner; solver dispatch, path registry, result batching, the pump, and diagnostics stay on the engine. The 9 ad-hocliquidity_verifier::*call sites inpy_binding.rscollapse ontoLiquidityMap::verify_against_onchain.Result batching stays on
ArbitrageEngine—result_tx,delivered, andcompute_diff_and_send(incrementalResultBatchdiffs) are solver-shaped (diffing the solver’sresultsagainst what Python hasdelivered), not state-shaped. BotCore has no role in result batching.The eager-processing architecture is preserved literally:
apply_logmutates BotCore immediately (no buffer, no lag),solve_dirtycoalesces the re-derive+solve exactly as today.Reorg rollback (
restore_before_block) reaches the live hot path for the first time — currently the pump applies events forward only and ignores theeth_subscriberemoved: boolflag on log events (the canonical reorg signal). Wiring this is a new behavior, not a mechanical port.RustPoolCache/ArbPoolCacheAdapterare retired by deletion (see “Legacy solver path retirement” below) — their third live copy dissolves along with the RustRustPoolCachePyO3 surface itself.Single-pool swap calc (
calculate_tokens_out/calculate_tokens_in) and swap encoding (encode_swap) stay onBotCore— per-pool math over state, mirroring Python’scalculate_tokens_out_from_tokens_in. The V3 stub needs implementation (V3/V4 single-pool CL swap math required for the future-state library consumer pattern).
Generalization discipline¶
LiquidityMap is implemented as concrete per-family types (LiquidityMap<V3PoolState>, LiquidityMap<V4PoolState>) sharing the generic LiquidityEventBuffer. No trait abstraction yet — only two CL families share the shape today, and Curve’s per-block-cache architecture is explicitly different (mirror-free, inline dependency resolution). A third sample is required before extracting an abstraction, per the project’s ruling against abstracting against a sample of one.