01The question
A single financial decision draws on evidence that has almost nothing in common: dense hourly price action, unstructured and frequently ironic narrative text, sparse scheduled macro releases, and heavy-tailed systemic risk. A monolithic model has two options, and neither is satisfactory. It can force everything into one shared representation and discard the structure that made each source informative, or it can disentangle the modalities internally, with no guarantee that a human can inspect how. Meanwhile, regulators increasingly require AI-driven financial recommendations to be explainable.
How should a heterogeneous financial decision be decomposed into modality-specialised agents — and how should their partial, unequally reliable outputs be recombined?
Our hypothesis: specialisation reduces information interference between modalities, and therefore improves robustness under regime shifts. Decomposition confines each modality’s noise to its own agent, and lets cross-modality influence enter through one explicit, inspectable fusion step.
Bitcoin is the case study, not the point. It has dense hourly data, a genuine narrative channel, explicit on-chain and geopolitical risk, and verifiable ground truth in realised returns: an unusually complete instance of the general problem.
02Architecture
Three independent specialists write typed outputs into shared state; a coordinator reads them and decides. The topology itself is a standard multi-agent pattern. The contribution is what travels along it.
LangGraph StateGraph · select an agent to see what it contributes
BUY / SELL / HOLD
with a natural-language rationale
Technical Analysis agent
- model
- Conv1D [64, 128, 256] → 2-layer LSTM (h = 256) → classification head
- input
- hourly BTC/USDT, 81 of 128 engineered features, 30-bar window
- stage 1
- trade / no-trade gate, threshold θ = 0.45
- stage 2
- long / short direction, plus a regression branch for exit magnitude
- result
- 66.5% direction accuracy on 5,385 test trades
Sentiment & Macro agent
- sources
- news 0.40 · social 0.30 · macro 0.30 (CPI prints, Fed decisions, tariffs)
- scoring
- FinBERT per item, s = p₊ − p₋, down-weighted after 24 hours
- reasoning
- Groq LLaMA-3-8b rereads hedged and forward-guidance text
- signal
- BUY if S > 0.3 · SELL if S < −0.3 · otherwise HOLD
- fallback
- keyword model when FinBERT is unavailable
Risk & Volatility agent
- sources
- on-chain 0.60 (volume, hash rate, mempool, addresses, miner outflows) · geopolitical 0.40
- model
- LightGBM over 62 volatility-oriented features
- label
- 48-bar forward drawdown: < 1% · 1–3% · > 3%
- signal
- LOW · MEDIUM · HIGH risk
- fallback
- heuristic over the same weights, then MEDIUM_RISK at confidence 0 — never an exception
Coordinator
- fusion
- D = 0.45·c₁v₁ + 0.35·c₂v₂ + 0.20·c₃v₃, where vᵢ ∈ {+1, 0, −1} and cᵢ is the agent’s own confidence
- decision
- BUY if D ≥ +0.20 · SELL if D ≤ −0.20 · otherwise HOLD
- veto
- HOLD if risk = HIGH and |D| < 0.55
- rationale
- an LLM merges the three justifications into signal, evidence and the main risk caveat
- runtime
- LangGraph StateGraph; a closed-form legacy mode evaluates the same rule without it
The Coordinator runs the technical model in-process and the other two agents as subprocesses emitting JSON, so each can depend on its own library stack without version conflicts. The fixed node order exists for deterministic tracing, not because of a data dependency.
03Contributions
The durable contributions are architectural, and none of them depends on the asset or the model classes used here.
-
01 · CONTRACT
A model-agnostic typed output
Every agent returns the same object: signal, continuous confidence, key factors, a natural-language justification and its data sources. A deep sequence model, an LLM-augmented classifier and a gradient-boosted tree are fused by one rule, with no per-agent integration code anywhere. Explainability is intrinsic, not a SHAP layer added after a black box has decided.
-
02 · RESILIENCE
Failures lower confidence, not uptime
Each specialist degrades through three tiers: trained model, heuristic fallback, safe default at confidence zero. The Coordinator always receives a well-formed output, so the fusion rule applies unconditionally. An expired API key becomes a down-weighted vote, not a crashed run.
-
03 · ROUTING
Uncertainty decides the weight
Each agent’s self-reported confidence and the assessed risk state decide how much its vote counts, including not at all. An uncertain agent is down-weighted automatically, and a discrete veto forces HOLD when risk is high and the fused signal is borderline.
Stated plainly: the fusion weights are fixed, hand-calibrated constants. That makes them transparent and auditable, and it also means the system does not adapt if an agent’s reliability drifts. Learning them from realised outcomes is the second item of future work.
04The specialists
Technical Analysis: seven years, split once
Hourly BTC/USDT candles from Binance, split in time with no overlap. From the raw series we over-generated 128 standard indicator features, from price action, momentum, volatility, volume and higher-timeframe context, then pruned them empirically rather than hand-picking a subset in advance.
Permutation importance shuffled each column and recorded the change in cumulative P&L; the 47 features whose removal improved it were dropped, leaving 81. The five strongest survivors are the 21-, 14- and 7-bar ATR, then price’s distance from its 9- and 50-period daily EMAs. The model leans hardest on how volatile the market is right now, and on where price sits inside its daily trend.
| Stage | Class | Precision | Recall | F1 |
|---|---|---|---|---|
| S1 · trade gate | no-trade | 0.996 | 0.128 | 0.226 |
| S1 · trade gate | trade | 0.622 | 1.000 | 0.767 |
| S2 · direction | short | 0.702 | 0.621 | 0.659 |
| S2 · direction | long | 0.633 | 0.712 | 0.670 |
Stage 1 accuracy 0.642 over 8,653 trades; stage 2 accuracy 0.665 over 5,385. Read stage 1 honestly: recall on no-trade is 0.128, so the gate admits almost every bar. It is a permissive filter, not a selective detector.
Sentiment & Macro: a classifier, then a reader
FinBERT scores each article and post off the shelf, deliberately so: a public, domain-adapted model with established provenance beats one we would have had to train and validate ourselves. Raw scores misread hedged and forward-guidance language, so an LLM then rereads the top snippets alongside the aggregate scores and writes a contextual interpretation. Only the two CNN–LSTM models and the LightGBM classifier were trained by us, and only because no off-the-shelf model exists for those tasks.
Risk & Volatility: separation that is real but modest
Over the test period, the share of bars followed by a severe drawdown rises with the predicted risk class:
Severe = more than 3% drawdown over the next 48 bars; the base rate is 25.7%. Mean 48-hour drawdown rises from 1.78% to 2.40% to 4.02%. As a binary high-risk detector, macro F1 is 0.550: the separation is real, and it is modest.
05Results
Every bar is assigned a triple-barrier return, and reported P&L is the arithmetic sum of those returns over overlapping positions, minus 0.1% per side. That is why buy-and-hold reads −127.7%: impossible for a real holding, an artefact of the convention. The magnitudes below are comparable with each other, and not with returns reported anywhere else.
Why a CNN–LSTM
Same 81 features, same split, same convention, four model families:
CNN–LSTM ranks highest and the non-temporal baseline is unprofitable, consistent with temporal context mattering on this feature set. All four share the same selected features, so the comparison is internally consistent and inherits the selection caveat below.
Testing the hypothesis: add one agent at a time
The result that matters is not P&L. Maximum drawdown does not improve when sentiment or risk is added alone. It improves only when the two act together:
Maximum drawdown, bars scaled to −80%. Shorter is better.
| Combination | Cum. P&L | Sharpe | Win rate | Max DD |
|---|---|---|---|---|
| Buy & hold | −127.7% | — | — | — |
| TA only | +1,799% | 18.87 | 65.7% | −71.9% |
| TA + sentiment | +1,612% | 17.91 | 65.3% | −71.2% |
| TA + risk | +1,726% | 18.20 | 65.1% | −71.9% |
| TA + sentiment + risk | +1,269% | 15.40 | 65.0% | −68.8% |
| Full system, with veto | +1,259% | 15.36 | 65.0% | −68.8% |
Two observations. P&L and Sharpe fall monotonically as agents are added, the expected cost of gating marginal trades through more, sometimes-disagreeing evidence. And because the single-agent rows bracket the joint row, the drawdown effect is not explained by either signal on its own: it arises at the coordinator’s fusion step, which is the form of evidence the hypothesis predicts. It is 3.1 points, on one asset, over one window. Suggestive, not established. The discrete veto adds little beyond continuous weighting.
06Scoring the reasoning
Backtest metrics say nothing about whether an agent reasoned correctly on any single inference. So we packaged live input–output pairs as JSON and had Gemini 2.5 Flash score each agent’s reasoning from 1 to 10, with three recommendations each, across four rounds of increasingly strict prompting.
| Agent | Score | Primary findings |
|---|---|---|
| Technical | 9 / 10 | Coherent signal and confidence; threshold logic correct. |
| Sentiment | 9 / 10 | Strong FinBERT integration; LLM reasoning aligned. |
| Risk | 6 / 10 | Noisy geopolitical events; empty volatility metrics. |
| Coordinator | 9 / 10 | Transparent fusion; each sub-agent’s contribution visible. |
We treat the scores as soft evidence: one model’s judgement, not a validated instrument. The durable finding is what the protocol surfaced. The risk agent had two silent defects. Its geopolitical pipeline admitted irrelevant articles, from entertainment to local sports, tagged as political instability, diluting every risk score it produced. And its volatility metrics came back empty despite a non-zero volatility score: a pipeline gap hidden behind a plausible number.
Neither defect raises an exception, changes the shape of an output, or moves aggregate P&L. Unit tests miss them because the code behaves as written. Backtests miss them because the metrics stay plausible.
Inspecting per-inference reasoning traces caught both. That is the result we consider most transferable: reasoning traces are a distinct observability surface for multi-agent systems, and we expect that to hold well beyond finance.
07Limitations
- Selection was not nested. Feature selection and both hyperparameter searches were scored on the test window, not validation, so every test figure is in-sample with respect to model selection. We would expect a properly nested protocol to produce materially lower figures.
- The accounting convention is non-standard. Summed rather than compounded returns over overlapping positions, without capital constraints or slippage, with correspondingly inflated Sharpe values.
- The sentiment check is not independent. Financial PhraseBank is FinBERT’s own fine-tuning corpus, so our 94.73% there only confirms the model loads and polarity is not inverted. It is not evidence of generalisation.
- The weights are hand-calibrated, and the evaluation covers one asset over one period. Domain-generality is an argument from construction, since no agent-specific logic sits in the coordinator, not an empirical result.
The architectural contributions do not depend on any of these trading figures. Rebuilding the results under a nested protocol with a compounded, non-overlapping return convention is the first item of future work.
08Cite
P. Canoski, B. Gjorgievski, M. Toshevska, S. Kalajdzhiski, and S. Gievska. “Adaptive Explainable Multi-Agent Intelligence for Heterogeneous Financial Decision Fusion: A LangGraph Framework with Dynamic Confidence-Aware Coordination.” Accepted at the 18th International Conference on ICT Innovations, Struga, North Macedonia, 26–28 September 2026.