The Green Screen Trap: Why a Trading Bot That Beat the Market Still Counted as a Failure
A developer let Claude and GPT trade crypto, beat the benchmark by 10 points, and promptly declared the run inconclusive. Here is why refusing to believe your own hype is the rarest engineering skill left.

The most dangerous moment for any software engineer or founder is when the dashboard turns bright green and you have no clean mathematical explanation for why.
That is the exact moment bad habits turn into fatal corporate beliefs. It is how tech founders raise seed rounds on vanity traffic spikes, how crypto bros mistake a bull market tide for trading genius, and how engineering teams push unvetted prompt pipelines into production because "it worked on my machine."
A developer named nunc just published a masterclass in intellectual self-defense. Over two months, they ran simulated, paper-trading crypto portfolios governed by LLMs on Base Layer-2 tokens. Season 1 was a mess: the bots churned fees, misread market beta, and lost to simple cash.
Season 2 looked like an unqualified triumph: all four LLM experimental arms beat the passive 50/50 benchmark by 3 to 10 percentage points. Claude Opus structured wrapped at +9.83 pp over benchmark. GPT structured delivered +8.87 pp.
Most people would take a screenshot, slap it on X, launch a Substack, and pitch an "autonomous AI hedge fund" to naive angels.
Instead, the author slapped down a verdict of inconclusive and killed the celebration. Why? Because before the code executed a single tick, they pre-registered strict validation criteria. And when tested against cost-sensitivity sweeps and double-window consistency, the green numbers cracked.
The interesting thing about this story is not merely that an LLM beat a passive crypto index over 30 days. It is actually how easily builders fool themselves into thinking stochastic text generators have strategic insight, when all they did was ride hidden asset beta wrapped in deterministic code guardrails.
The Architecture: Keeping the Model in a Cage
If you take nothing else away from this experiment, take the system architecture. Most "AI agent" tutorials floating around developer circles are architectural malpractice: they wire an API key to a model, hand it tool-calling privileges, and pray the model doesn't blow through bank balances on token hallucinations.
Nunc did the opposite. The architecture treated the LLM as an untrusted, high-variance advisory engine trapped behind an uncompromising Python policy firewall:
- Zero Secret Leakage: The models were called via vendor CLIs (
claude -p,codex exec) under sandboxed user sessions without environment variables. - Deterministic Hard Limits: Trade limits ($5 to $20), frequency caps (maximum 3 trades per day), portfolio meme-token exposure limits (capped at 15%), and an unyielding kill switch (liquidate everything to USDC if baseline dropped below 70%) lived entirely in Python code. The prompt had zero authority over these bounds.
- Structured vs. Autonomous: The core experiment tested whether the model should operate autonomously (staring at raw data and deciding) or through a structured pipeline (Python code ranks candidate tokens, generates a cost table, forces a daily forecast, and enforces an edge gate).
The result on this axis was definitive: Code structure slaughtered model autonomy.
Opus Structured beat Opus Autonomous by +3.24 pp. GPT Structured beat GPT Autonomous by +5.49 pp.
When you strip away the romantic sci-fi fantasy of an "autonomous agent," an LLM is essentially an unpredictable intern prone to over-trading and narrative rationalization. When you let code constrain its search space and force it to justify edges against explicit math, performance jumps. But even that jump hides a structural trap.
The Reality of Hidden Friction: The Nigerian Market Analogy
To understand why the author refused to call Season 2 a victory, you have to look at how real commerce works when paper theories meet friction.
Think about an importer moving physical inventory from the Onitsha Main Market down to retail shops in Owerri. On paper, the margin looks sweet: buy cheap at bulk wholesale, mark it up 18% at retail, project a clean profit.
Then reality intervenes: diesel prices spike on the expressway, local council agberos demand unpredictable haulage fees, a minor delay at the transit park leaves boxes water-damaged, and by the time money lands in your hand, your 18% margin turned into a 3% net loss. You worked for the transport union, not yourself.
That is exactly what happened to the GPT trading agent. On the surface, GPT structured posted an $8.87 margin above passive hold. But the author’s pre-registered rules required a cost-sensitivity sweep at 0.5×, 1×, and 2× variable costs (factoring in slippage haircuts and gas costs).
At 2× costs, GPT’s edge collapsed. The agent was churning trades where the theoretical margin evaporated the moment simulated market liquidity tightened. Claude Opus held up better under the cost sweep, but failed the multi-window reproducibility standard.
The green number on the terminal was an artifact of favorable timing and low simulated slippage, not a durable machine learning moat.
The Short Answer
LLMs do not possess native market alpha. What looks like trading intelligence in an LLM agent is almost always market beta, luck, or the deterministic genius of the Python wrappers boxing the model in. If your agentic system cannot survive pre-registered friction tests and cost sweeps, you do not have an AI product; you have an expensive random number generator with a high-bandwidth vocabulary.
What Is Really Happening
The AI developer ecosystem is drunk on output and starved of methodology.
In Gbagada workstations, Silicon Valley garages, and London hackathons, founders are running ad-hoc evals. They tweak a prompt, watch their terminal show a favorable metric on a single test run, call it "product-market fit," and deploy.
What this experiment exposes is the massive delta between evaluation theater and scientific pre-registration:
- In Season 1, the author caught themselves: without a frozen decision rule, they would have looked at Opus finishing at $101.48 and shouted, "We're profitable, ship it!" But that profit was just early-cycle ETH beta.
- In Season 2, the author enforced scientific hygiene: fixing the model version, universe, and prompt before tick 1, and deciding what constituted a "win" before looking at the data.
When builders skip pre-registration, they fall prey to post-hoc curve fitting. They rewrite the narrative after the trade executes to convince themselves the machine was clever all along.
The Assumption I'd Challenge
The assumption: An LLM agent can identify subtle financial or operational signals that deterministic algorithms miss.
The challenge: High confidence: The model is not finding an edge. It is acting as a semantic summarizer of current sentiment, which means it is permanently lagging or tracking momentum beta. In this experiment, the moment you removed the code-side edge gate and the code-ranked candidate table, the autonomous model’s performance degraded immediately.
The value was never in the model’s "reasoning." The value was in the Python code that filtered out garbage before the model was even allowed to speak, and the strict risk gates that blocked the model from self-destructing. The LLM was simply the costliest component in a system that succeeded because of classic software engineering.
The Strategic Options
If you are a builder looking to integrate autonomous LLM decision-making into mission-critical workflows (finance, logistics dispatch, dynamic pricing, inventory reordering), you have three choices:
| Path | Architectural Reality | Economic & Operational Trade-off |
|---|---|---|
| Option A: Pure Autonomy | Hand raw context to the model, give it tool access, let it act on live endpoints. | High ruin risk. High token costs, rapid fee churn, zero auditability, high hallucination risk when market regimes switch. |
| Option B: Deterministic Sandboxing (The "nunc" Pattern) | Code ranks candidates, computes cost floors, forces mathematical outputs; model only selects from code-approved actions. | Moderate build complexity, high safety. You treat the model as an advisory heuristic, but the Python kernel enforces all invariant risk rules. |
| Option C: Pure Deterministic Heuristics (Drop the LLM) | Replace the LLM entirely with classical rule-based optimization, statistical regression, or simple cron-gated logic. | Lowest operating cost, zero latency. You lose the semantic adaptability of language models, but eliminate API failure modes and token overhead. |
My Recommendation
If you are building production systems, choose Option B, but constantly audit whether you can downgrade to Option C.
Do not give an LLM execution privileges. Ever. Build your architecture so that the policy engine fails closed. If your system loses internet access, if the model hallucinates a decimal place, or if vendor rate limits throttle your CLI calls, the system should default to cash or passive holds—not panic-execute.
More importantly: adopt pre-registration in your internal product evaluation. Before you test an AI agent against a business problem, write down in a markdown file:
- What exact metric proves superiority over a boring script?
- What happens to the math when costs (APIs, customer support, platform cuts) double?
- How many distinct time windows must this survive before we put customer money or data behind it?
If you don't write down the rules before you start the run, your brain will rationalize whatever green numbers spit out at the end.
What I Would Do Next
- Implement Hard Negative Tests: Take your existing agent prompts and run them against historical periods where the market or your business metrics tanked. If the bot didn't liquidate or sit on its hands, your risk gates are broken.
- Stress-Test Variable Costs: Recalculate your system’s unit economics assuming OpenAI, Anthropic, or infrastructure fees double, and execution slippage increases by 100%. If your margin vanishes under friction, you have an illusion, not a business model.
- Audit Token Spend vs. Value Add: Measure the raw token cost spent on Claude Opus or GPT-4 reasoning against the incremental gain over a deterministic 7-day moving average. If the delta is less than the compute bill, strip the LLM out of the loop.
What Would Change My Mind
I will change my mind and back autonomous LLM-led execution systems if:
- An LLM trading architecture demonstrates a positive alpha (>3% net of real execution slippage, gas, and API compute costs) across four consecutive quarters across both bull and bear market regimes.
- The system achieves this without a deterministic human-coded policy engine telling it what tokens it is allowed to evaluate or capping its position sizes.
Until someone shows that proof on a pre-registered ledger, keep your model on a tight leash, let Python run the bank, and never confuse a green day on a paper dashboard with genuine operational edge.
Related from Engineering
Let's build your next big product.
Accepting project-based freelance, remote engineering roles, and hybrid positions.