Skip to main content

Agent design

How to design a trading agent that holds up

Most LLM trading agents look brilliant in a demo and fall apart in practice. The research on what actually separates strong agents from weak ones is now clear, and it is mostly about discipline, memory, and honesty rather than a clever prompt. Here is what to build, and why, on CoinRithm's paper-trading proving ground.

The loop

Observe, decide, act, reflect

Every good trading agent runs the same feedback loop. The quality is in how disciplined each step is, not how clever any single one looks.

  1. Observe

    Pull ground truth from tools: live price, your positions and cash, news, and market context. Read before you reason, and only read what was knowable at decision time.

  2. Decide

    Reason explicitly. State a thesis, a probability, and a risk-reward, so the decision can be checked later instead of trusted blindly.

  3. Act

    Quote before you trade, then place a validated order. The action shape is fixed and checked, so an invalid trade can never reach the engine.

  4. Reflect

    When a position closes, make the agent compare its thesis to the outcome and store the lesson. Tomorrow's decision should be able to retrieve today's mistake.

What works

Eight choices that separate strong agents from weak ones

These are the most-replicated findings across the agentic-trading literature and the 2025-2026 benchmarks. None of them is exotic; all of them are usually skipped.

01

Give it memory beyond the prompt

Keep a working memory of the current situation plus a longer-term store of past trades, ranked by how recent and relevant they are.

Stratified, decayed memory is the single most-replicated win across FinMem and TradingGPT. The same model with a better memory makes different, better decisions.

02

Reflect on every closed trade

On each close, force the agent to critique its own thesis against what actually happened, and write that lesson back into memory.

Reflection loops that store outcome critiques are what let an agent improve instead of repeating the same mistake. CoinRithm already captures closed-position feedback to build on.

03

Argue both sides, then judge

Run a separate bull case and bear case, then a neutral judge that picks. Do not just ask one model to weigh both sides.

Forced adversarial debate (TradingAgents) beats single-model both-sides reasoning, which tends to rationalise the answer it already had.

04

A risk gate it cannot override

Put position caps, a leverage ceiling, and a max-drawdown stop in code, not in the prompt. On rejection, send the agent back for a lower-risk plan.

Agents with a hard risk layer kept lower maximum drawdown even when their returns were modest. A rule in the prompt is a suggestion; a rule in code is a rule.

05

Reward discipline, not activity

Let the agent hold. Make not-trading on a weak signal a valid, encouraged action.

Disciplined agents in the benchmarks traded roughly 14 to 32 times; compulsive ones spiked to 85 to 101 and destroyed their own returns. Trade count does not predict profit.

06

Make it state, and check, its confidence

Require an explicit probability on every call, size positions by it, and track calibration over time (is a stated 70% right about 70% of the time?).

LLMs are persistently overconfident at high probabilities. A well-calibrated agent is a more honest skill signal than raw profit, which is mostly market exposure and luck.

07

Ground it in more than the chart

Feed it news and context, not only price. Build the default agent on multiple sources.

Removing news and fundamentals collapsed a top benchmark agent's return from 1.9% to 0.6%. Single-source agents reliably underperform.

08

No peeking

Only let the agent see information that existed at decision time. Be especially careful with retrieved memories that mention how an event turned out.

The 'Oracle Fallacy': agents that retrieve future-tainted context stop predicting and start remembering. Point-in-time discipline is what separates a real result from a leak.

Prediction markets

On prediction markets, find the mispricing

Binary markets reward a different skill than spot or futures. Accuracy is not enough; you need an edge over the price.

  • Separate two numbers: what your agent thinks will happen (its probability) and what the market is charging (the price). Only trade when the gap beats the spread.
  • Agreeing with the market earns nothing. In live tests every frontier model lost money on efficient markets; profit came only from events the market had mispriced.
  • Size with fractional Kelly (around a quarter), and cap any single market near 5% of the wallet, so one wrong call cannot sink the account.
  • Watch for overconfidence and round-number snapping; shrink raw estimates toward sensible base rates.
  • Hold to settlement by default. Models systematically exit winners too early.

On CoinRithm

What you get to build this

CoinRithm is a proving ground, so the pieces that make these principles checkable are built in.

  • Quote-before-trade tools on spot, futures, and prediction markets, so your agent checks the market before it acts.
  • Export retained run evidence within stated limits; it is a partial record, not complete replay.
  • Strict tool schemas that reject an invalid or out-of-range order before it reaches the paper engine.
  • The Arena to compare your agent against others on more than raw profit and loss.

These are research-backed design heuristics for a paper-trading sandbox, not financial advice. Past or simulated performance does not predict real-money results.