Home/Research

18th International Conference on ICT InnovationsStruga · 26–28 September 2026

Adaptive Explainable Multi-Agent Intelligence for Heterogeneous Financial Decision Fusion

A LangGraph framework with dynamic, confidence-aware coordination.

Petar Canoski · first and corresponding author, Borjan Gjorgievski, Martina Toshevska, Slobodan Kalajdzhiski, Sonja Gievska

Faculty of Computer Science and Engineering, Ss. Cyril and Methodius University in Skopje

Stage-2 direction
66.5%
accuracy on 5,385 test trades
Risk classifier
0.550
macro F1, high-risk vs. severe drawdown
Data
~61k
hourly bars, Mar 2019 – Feb 2026
Risk veto
4.6%
of test bars forced to HOLD (423 of 9,133)

01The question

A single financial decision draws on evidence that has almost nothing in common: dense hourly price action, unstructured and frequently ironic narrative text, sparse scheduled macro releases, and heavy-tailed systemic risk. A monolithic model has two options, and neither is satisfactory. It can force everything into one shared representation and discard the structure that made each source informative, or it can disentangle the modalities internally, with no guarantee that a human can inspect how. Meanwhile, regulators increasingly require AI-driven financial recommendations to be explainable.

How should a heterogeneous financial decision be decomposed into modality-specialised agents — and how should their partial, unequally reliable outputs be recombined?

Our hypothesis: specialisation reduces information interference between modalities, and therefore improves robustness under regime shifts. Decomposition confines each modality’s noise to its own agent, and lets cross-modality influence enter through one explicit, inspectable fusion step.

Bitcoin is the case study, not the point. It has dense hourly data, a genuine narrative channel, explicit on-chain and geopolitical risk, and verifiable ground truth in realised returns: an unusually complete instance of the general problem.

02Architecture

Three independent specialists write typed outputs into shared state; a coordinator reads them and decides. The topology itself is a standard multi-agent pattern. The contribution is what travels along it.

LangGraph StateGraph · select an agent to see what it contributes

BUY / SELL / HOLD
with a natural-language rationale

Technical Analysis agent

model
Conv1D [64, 128, 256] → 2-layer LSTM (h = 256) → classification head
input
hourly BTC/USDT, 81 of 128 engineered features, 30-bar window
stage 1
trade / no-trade gate, threshold θ = 0.45
stage 2
long / short direction, plus a regression branch for exit magnitude
result
66.5% direction accuracy on 5,385 test trades

Sentiment & Macro agent

sources
news 0.40 · social 0.30 · macro 0.30 (CPI prints, Fed decisions, tariffs)
scoring
FinBERT per item, s = p₊ − p₋, down-weighted after 24 hours
reasoning
Groq LLaMA-3-8b rereads hedged and forward-guidance text
signal
BUY if S > 0.3 · SELL if S < −0.3 · otherwise HOLD
fallback
keyword model when FinBERT is unavailable

Risk & Volatility agent

sources
on-chain 0.60 (volume, hash rate, mempool, addresses, miner outflows) · geopolitical 0.40
model
LightGBM over 62 volatility-oriented features
label
48-bar forward drawdown: < 1% · 1–3% · > 3%
signal
LOW · MEDIUM · HIGH risk
fallback
heuristic over the same weights, then MEDIUM_RISK at confidence 0 — never an exception

Coordinator

fusion
D = 0.45·c₁v₁ + 0.35·c₂v₂ + 0.20·c₃v₃, where vᵢ ∈ {+1, 0, −1} and cᵢ is the agent’s own confidence
decision
BUY if D ≥ +0.20 · SELL if D ≤ −0.20 · otherwise HOLD
veto
HOLD if risk = HIGH and |D| < 0.55
rationale
an LLM merges the three justifications into signal, evidence and the main risk caveat
runtime
LangGraph StateGraph; a closed-form legacy mode evaluates the same rule without it

The Coordinator runs the technical model in-process and the other two agents as subprocesses emitting JSON, so each can depend on its own library stack without version conflicts. The fixed node order exists for deterministic tracing, not because of a data dependency.

03Contributions

The durable contributions are architectural, and none of them depends on the asset or the model classes used here.

  • 01 · CONTRACT

    A model-agnostic typed output

    Every agent returns the same object: signal, continuous confidence, key factors, a natural-language justification and its data sources. A deep sequence model, an LLM-augmented classifier and a gradient-boosted tree are fused by one rule, with no per-agent integration code anywhere. Explainability is intrinsic, not a SHAP layer added after a black box has decided.

  • 02 · RESILIENCE

    Failures lower confidence, not uptime

    Each specialist degrades through three tiers: trained model, heuristic fallback, safe default at confidence zero. The Coordinator always receives a well-formed output, so the fusion rule applies unconditionally. An expired API key becomes a down-weighted vote, not a crashed run.

  • 03 · ROUTING

    Uncertainty decides the weight

    Each agent’s self-reported confidence and the assessed risk state decide how much its vote counts, including not at all. An uncertain agent is down-weighted automatically, and a discrete veto forces HOLD when risk is high and the fused signal is borderline.

D = 0.45c₁v₁ + 0.35c₂v₂ + 0.20c₃v₃   BUY if D ≥ +0.20 · SELL if D ≤ −0.20
HOLD if risk = HIGH and |D| < 0.55   the discrete risk veto

Stated plainly: the fusion weights are fixed, hand-calibrated constants. That makes them transparent and auditable, and it also means the system does not adapt if an agent’s reliability drifts. Learning them from realised outcomes is the second item of future work.

04The specialists

Technical Analysis: seven years, split once

Hourly BTC/USDT candles from Binance, split in time with no overlap. From the raw series we over-generated 128 standard indicator features, from price action, momentum, volatility, volume and higher-timeframe context, then pruned them empirically rather than hand-picking a subset in advance.

2019-03-02 → 2024-01-19training 2024-01-19 → 2025-02-04validation 2025-02-05 → 2026-02-20test

Permutation importance shuffled each column and recorded the change in cumulative P&L; the 47 features whose removal improved it were dropped, leaving 81. The five strongest survivors are the 21-, 14- and 7-bar ATR, then price’s distance from its 9- and 50-period daily EMAs. The model leans hardest on how volatile the market is right now, and on where price sits inside its daily trend.

Technical agent classification metrics on the test set
StageClassPrecisionRecallF1
S1 · trade gateno-trade0.9960.1280.226
S1 · trade gatetrade0.6221.0000.767
S2 · directionshort0.7020.6210.659
S2 · directionlong0.6330.7120.670

Stage 1 accuracy 0.642 over 8,653 trades; stage 2 accuracy 0.665 over 5,385. Read stage 1 honestly: recall on no-trade is 0.128, so the gate admits almost every bar. It is a permissive filter, not a selective detector.

Sentiment & Macro: a classifier, then a reader

FinBERT scores each article and post off the shelf, deliberately so: a public, domain-adapted model with established provenance beats one we would have had to train and validate ourselves. Raw scores misread hedged and forward-guidance language, so an LLM then rereads the top snippets alongside the aggregate scores and writes a contextual interpretation. Only the two CNN–LSTM models and the LightGBM classifier were trained by us, and only because no off-the-shelf model exists for those tasks.

Risk & Volatility: separation that is real but modest

Over the test period, the share of bars followed by a severe drawdown rises with the predicted risk class:

Severe = more than 3% drawdown over the next 48 bars; the base rate is 25.7%. Mean 48-hour drawdown rises from 1.78% to 2.40% to 4.02%. As a binary high-risk detector, macro F1 is 0.550: the separation is real, and it is modest.

05Results

Before any number

Every bar is assigned a triple-barrier return, and reported P&L is the arithmetic sum of those returns over overlapping positions, minus 0.1% per side. That is why buy-and-hold reads −127.7%: impossible for a real holding, an artefact of the convention. The magnitudes below are comparable with each other, and not with returns reported anywhere else.

Why a CNN–LSTM

Same 81 features, same split, same convention, four model families:

CNN–LSTM ranks highest and the non-temporal baseline is unprofitable, consistent with temporal context mattering on this feature set. All four share the same selected features, so the comparison is internally consistent and inherits the selection caveat below.

Testing the hypothesis: add one agent at a time

The result that matters is not P&L. Maximum drawdown does not improve when sentiment or risk is added alone. It improves only when the two act together:

Maximum drawdown, bars scaled to −80%. Shorter is better.

Incremental agent-combination ablation
CombinationCum. P&LSharpeWin rateMax DD
Buy & hold−127.7%———
TA only+1,799%18.8765.7%−71.9%
TA + sentiment+1,612%17.9165.3%−71.2%
TA + risk+1,726%18.2065.1%−71.9%
TA + sentiment + risk+1,269%15.4065.0%−68.8%
Full system, with veto+1,259%15.3665.0%−68.8%

Two observations. P&L and Sharpe fall monotonically as agents are added, the expected cost of gating marginal trades through more, sometimes-disagreeing evidence. And because the single-agent rows bracket the joint row, the drawdown effect is not explained by either signal on its own: it arises at the coordinator’s fusion step, which is the form of evidence the hypothesis predicts. It is 3.1 points, on one asset, over one window. Suggestive, not established. The discrete veto adds little beyond continuous weighting.

06Scoring the reasoning

Backtest metrics say nothing about whether an agent reasoned correctly on any single inference. So we packaged live input–output pairs as JSON and had Gemini 2.5 Flash score each agent’s reasoning from 1 to 10, with three recommendations each, across four rounds of increasingly strict prompting.

LLM-as-a-Judge evaluation, final prompt revision
AgentScorePrimary findings
Technical9 / 10Coherent signal and confidence; threshold logic correct.
Sentiment9 / 10Strong FinBERT integration; LLM reasoning aligned.
Risk6 / 10Noisy geopolitical events; empty volatility metrics.
Coordinator9 / 10Transparent fusion; each sub-agent’s contribution visible.

We treat the scores as soft evidence: one model’s judgement, not a validated instrument. The durable finding is what the protocol surfaced. The risk agent had two silent defects. Its geopolitical pipeline admitted irrelevant articles, from entertainment to local sports, tagged as political instability, diluting every risk score it produced. And its volatility metrics came back empty despite a non-zero volatility score: a pipeline gap hidden behind a plausible number.

Neither defect raises an exception, changes the shape of an output, or moves aggregate P&L. Unit tests miss them because the code behaves as written. Backtests miss them because the metrics stay plausible.

Inspecting per-inference reasoning traces caught both. That is the result we consider most transferable: reasoning traces are a distinct observability surface for multi-agent systems, and we expect that to hold well beyond finance.

07Limitations

What these numbers cannot establish
  1. Selection was not nested. Feature selection and both hyperparameter searches were scored on the test window, not validation, so every test figure is in-sample with respect to model selection. We would expect a properly nested protocol to produce materially lower figures.
  2. The accounting convention is non-standard. Summed rather than compounded returns over overlapping positions, without capital constraints or slippage, with correspondingly inflated Sharpe values.
  3. The sentiment check is not independent. Financial PhraseBank is FinBERT’s own fine-tuning corpus, so our 94.73% there only confirms the model loads and polarity is not inverted. It is not evidence of generalisation.
  4. The weights are hand-calibrated, and the evaluation covers one asset over one period. Domain-generality is an argument from construction, since no agent-specific logic sits in the coordinator, not an empirical result.

The architectural contributions do not depend on any of these trading figures. Rebuilding the results under a nested protocol with a compounded, non-overlapping return convention is the first item of future work.

08Cite

P. Canoski, B. Gjorgievski, M. Toshevska, S. Kalajdzhiski, and S. Gievska. “Adaptive Explainable Multi-Agent Intelligence for Heterogeneous Financial Decision Fusion: A LangGraph Framework with Dynamic Confidence-Aware Coordination.” Accepted at the 18th International Conference on ICT Innovations, Struga, North Macedonia, 26–28 September 2026.

Author’s version (PDF)