Meta Description: Verified 2026 case studies: Claude Portfolio +19.04% vs 12.24% S&P with $27M copy capital, HKU Agentic Trader Qwen +9.9% to DeepSeek -15.1%, NBER factor trap, TradingAgents 53k stars. Complete AI-augmented workflow for swing traders.
URL Slug: ai-llm-swing-trading-retail-case-studies-2026
Primary Keyword: AI swing trading retail investors
Related Keywords: ChatGPT trading bot, Claude trading portfolio Autopilot, HKU Agentic Trader, LLM stock picking, copy trading AI, TradingAgents framework, AI investment committee, Lopez-Lira AI portfolio, retail AI hedge fund, agentic trading, swing trading catalyst, AI risk management
AI-Powered Trading by Retail Investors: What Recent Case Studies Actually Show About LLM Swing Trading Profits in 2026
PART 1 OF 6-10 The shift no one expected: The most profitable retail use of LLMs in 2026 isn't "Ask ChatGPT to pick a winner." It's using LLMs as a 24/7 research department, risk veto, and portfolio assistant. This series dissects live-money experiments with public capital — not backtests.
This is a polished analysis of how retail investors have effectively used AI large language models to achieve profits through swing trading and other strategies — with wins, losses, and the workflows that actually work.
Why This Topic Matters Right Now
For years, retail investors were told they were "dumb money" — chasing rallies, panic-selling dips, lacking tools. In 2026, that narrative cracked. Three things converged:
- Live copy capital: AI-managed portfolios on Autopilot attracted ~$27M into a single Claude portfolio after +19.04% vs 12.24% S&P 500 by Aug 4, and ~$200M across seven AI portfolios with 52,000 investors by July.
- Live agent benchmarks: The University of Hong Kong's Artificial Intelligence Evaluation Lab launched Agentic Trader — 10 LLMs trading live FX with identical $100k, identical tools, identical leverage. Qwen produced ~+9.9% to +10%, DeepSeek lost -15.1% over six weeks.
- Factor reality check: An NBER working paper found LLM portfolios tilted to large-cap, high-momentum, high-beta, low book-to-market — alpha disappeared after adjusting for risk factors.
Bottom line: The edge is real, but it's not magic stock picking. It's information compression, discipline, and risk control.
March to Aug 4, $50k seed
Table of Contents — Full 12,000 Word Series
- Part 1 (This Article): Introduction, why it matters, foundational concepts, and Claude/Autopilot real-money case study
- Part 2: HKU Agentic Trader Deep Dive — Why Qwen won, why DeepSeek lost, risk preferences, leverage traps
- Part 3: NBER Factor Critique — Is it alpha or beta? How to benchmark LLM portfolios correctly
- Part 4: TradingAgents Open-Source Hedge Fund — 7-agent architecture, bull/bear debate loop, how retail runs it locally on Ollama/Qwen3
- Part 5: The Retail Workflow That Works — Screen → Catalyst → Multi-model Debate → Technical → Fundamental → Risk → Human Approval
- Part 6: Catalyst-Driven Swing Trading — Earnings, sector rotation, news-reversal with LLM prompts
- Part 7: Technical Analysis + Structured Data — How to feed LLMs RSI, ATR, volume, short interest without hallucination
- Part 8: Risk Management Breakthrough — Position sizing, max exposure, trade journaling with AI
- Part 9: YouTube Lab — 24 videos reviewed, building trading bots with ChatGPT, Claude Cowork, n8n
- Part 10: Future & Checklist — Investment committee model, FINRA warnings, and sustainable edge
Foundational Concepts You Need
What is Swing Trading?
Swing trading holds positions from 2 days to 3-6 weeks to capture a "swing" caused by a catalyst — earnings surprise, guidance change, new contract, regulatory shift, or sector momentum. Unlike day trading, it depends on identifying a temporary shift in probability, not minute-by-minute prediction.
What is an LLM in Trading?
Large Language Models (ChatGPT, Claude, Gemini, Grok, DeepSeek, Qwen, Kimi) are not price predictors. Their strength is synthesizing unstructured information — 10-Ks, 10-Qs, earnings transcripts (avg ~7,000 words), SEC filings, analyst revisions, news — into a structured thesis faster than manual reading.
What is Agentic Trading?
Agentic trading is the step beyond chatbot. Agents receive live market data, tool access (search, brokerage APIs), and autonomy to act — they interpret changing environments, manage risk, and face consequences of past decisions. HKU's Agentic Trader is the cleanest test: identical capital and conditions, only model capability differs.
Part 1 Deep Dive: The Strongest Evidence — Claude Portfolio with Real Retail Money
One of the most important developments in 2026 is that AI portfolios stopped being screenshots and became copyable products.
The setup: Finance professor Alejandro López-Lira (University of Florida) and AI Finance Labs run public experiments where models receive updated market information, company news and fundamentals, then make portfolio decisions. Portfolios were offered through Autopilot, allowing users to copy trades with positions publicly disclosed. Critically, Claude did not independently manage the portfolio. It was used for research and investment ideas, while Lopez-Lira and his team set the strategy and risk parameters and controlled the investment process. This human oversight distinction matters enormously.
Verified results:
- Claude Portfolio launched March with $50,000 seed. By August 4, +19.04% vs 12.24% S&P 500.
- Attracted ~$27M from investors copying trades.
- Seven AI portfolios (ChatGPT, Grok, DeepSeek, Claude etc.) collectively attracted ~52,000 investors and ~$200M by July.
- DeepSeek portfolio reported 76% vs 27% S&P in May, later Autopilot data showed 89.1% gain since Feb 2025 launch with ~$49.6M assets. Claude continued to 31.1% gain since March launch with ~$49.4M assets as of Sept 4.
Why is this interesting for swing traders? The workflow is: Market scanner → LLM catalyst analysis → multi-model debate → technical confirmation → fundamental verification → risk calculation → position sizing → human approval → automated monitoring → journal → AI review. That turns vague advice into a system.
What Successful Retail Workflows Have in Common
Across positive experiments, five patterns repeat:
- Current information: Model needs fresh data. Historical knowledge alone fails for swing horizons.
- Imposed process: Not "What stock to buy?" but broken tasks: screen, investigate, rank catalysts, thesis, risks, entry/exit, monitor, reassess.
- Multi-source fusion: Fundamentals + price + volume + news + macro + sentiment + risk. No single source wins.
- Constrained risk: HKU showed greater activity and risk-taking did not guarantee better returns. Taking on greater risk did not necessarily lead to higher returns.
- Continuous reassessment: Swing thesis is temporary. Once catalyst changes, trade should change.
Comparison Table: HKU Live Results Snapshot
| Model | 6-Week Cumulative Return | Trade Count | Risk Behavior |
|---|---|---|---|
| Qwen (Alibaba) | +9.9% to +10% (strongest) | 500-800 | Managed risk carefully |
| Kimi / Seed | Among strongest | 500-800 | Balanced |
| GPT (OpenAI) | Near break-even | Varied | Cautious, low exposure |
| GLM | Near break-even | ~500-800 | Moderate |
| DeepSeek | -15.1% (largest loss) | 1,000+ | High leverage, larger drawdowns |
| Claude / Gemini | Substantial losses in HKU FX | 1,000+ each | Gemini high leverage |
Source: HKU Business School Agentic Trader report, April 2026 live evaluation, identical $100k, identical tools. Range varied significantly e.g. 9.9% for Qwen to -15.1% for Deepseek.
Video Lab: Watch These Workflows in Action
For Part 1, we embed 4 core videos that demonstrate the exact shift from chatbot to investment team:
1. Agentic Trading Explained — What Changed
2. Claude Finance Agents — Full Install for Retail Research
3. Someone Open-Sourced a Hedge Fund — TradingAgents Framework
4. I Built a Trading Bot with ChatGPT — $2000 Live Test
Original Analysis: What Happened, Why It Matters, Who Benefits, Risks, What's Next
What happened: Retail brokerages connected AI agents to live accounts in H1 2026 (Finance Magnates Intelligence reported at least 10 brokers, Claude in 9 deployments). Investors stopped asking ChatGPT for picks and started copying audited AI portfolios with human risk overlay.
Why it matters: It removes the translation gap — how to turn a conversation into a consistent system. Copy trading + AI research = retail can access portfolio construction that used to require quant team.
Who benefits: Retail investors who systematize (not those seeking hot tips), platforms with transparent performance (Autopilot), and traders who use LLMs to reduce research time, emotional trading, and position-sizing mistakes.
Risks: Short track records (Claude <6 months when first outperformance claimed), factor exposure mistaken for intelligence, leverage and overtrading (HKU: high activity ≠ better), hallucinations, and FINRA warning: be wary of claims that AI can guarantee amazing investment returns and that AI can generate false information.
What could happen next: The "AI investment committee" model — ChatGPT fundamental thesis, Claude adversarial risk review, Grok sentiment, DeepSeek quantitative reasoning — attacking same trade from different angles. Competitive edge shifts from picking stocks to designing, constraining, and supervising AI systems.
FAQ — Part 1
Did Claude really make 19% in a few months?
Yes, verified via TradingView / Finance Magnates syndication: Claude Portfolio launched March with $50k, reported 19.04% by Aug 4 vs 12.24% S&P, with ~$27M copy capital. However, it was not fully autonomous — Lopez-Lira team set strategy and risk parameters. Track record was under 6 months, too early to assess sustainability.
Why did Qwen beat DeepSeek in HKU live trading if DeepSeek is strong at reasoning?
HKU explicitly concluded strong performance on reasoning, coding, knowledge benchmarks does not necessarily translate to good financial-market performance. Trading tests risk control, position management, sustained decision-making under uncertainty. Qwen managed risk carefully; DeepSeek used high leverage and had larger drawdowns.
Can I just copy the AI portfolios?
You can via Autopilot-style copy trading, but entry timing, execution prices, fees, taxes, slippage, portfolio size affect results. Paper returns ≠ your returns. Always check independently and speak to regulated adviser.
Transition to Part 2: Now that we've established the strongest real-money evidence — Claude/Autopilot's $27M copy phenomenon with human oversight — we need to stress-test the opposite extreme: fully autonomous agents with no strategy prescribed. In Part 2, we dissect HKU Agentic Trader trade-by-trade: why Qwen's cautious 500-800 trades beat 1,000+ trade bots, why high leverage failed, and how to build your own risk-managed agent that trades less but better.
[Part 1 Complete. Say "Go" or "Proceed" to generate Part 2.]
Part 2: HKU Agentic Trader — Why Qwen Gained ~+10% While DeepSeek Lost -15.1% in Live Trading
PART 2 OF 10 Live markets are the ultimate exam. In Part 1 we saw copy capital follow Claude +19.04% vs 12.24% S&P. In Part 2 we flip to fully autonomous agents with no human risk overlay — 10 LLMs, $100k each, same data, same leverage, 6 weeks live FX. Result: +9.9% to -15.1% spread. More trades did NOT mean more profit.
This deep dive explains why — and what swing traders can steal from the winners.
Recap: From Copy Trading to Autonomous Trading
Part 1 established the human-in-the-loop model: Lopez-Lira's team used Claude for research, set risk parameters, and let retail copy via Autopilot. It worked because human veto prevented hallucinated trades.
HKU's Artificial Intelligence Evaluation Lab (AIEL), led by Prof. Jack Jiang, asked the harder question: What if we remove the human and let 10 leading LLMs — GPT, Claude, Gemini, DeepSeek, Qwen, Grok, GLM, Kimi, MiniMax, Seed — trade autonomously in live foreign exchange (EUR/USD, GBP/USD, USD/JPY, S&P Index, precious metals) with identical starting capital, tool access, and leverage? No strategy prescribed. All decisions generated autonomously.
Methodology: The Cleanest Test Yet
HKU's report outlines:
- Participating Models: US and China state-of-the-art — OpenAI, Anthropic, Google, Alibaba (Qwen), Moonshot (Kimi), Zhipu (GLM), ByteDance (Seed), DeepSeek, xAI (Grok), MiniMax.
- Environment: Live FX market data, real spreads, slippage, time pressure. Unlike static benchmarks (math, coding), trading tests continuous interpretation and adaptation.
- What was measured: Timely decisions, active risk control, position management, sustained decision-making, not just correct static answer.
"Performance gaps between LLMs become increasingly pronounced when they are required to continuously interpret changing environments and adapt their strategies in real time." — Prof. Jack Jiang, HKU Business School
Results: Who Won, Who Lost, Why
1. Qwen Generated Strongest Cumulative Returns
By end of evaluation: Qwen ~+9.9% to +10%, Kimi and Seed among strongest, GLM and GPT broadly near break-even, MiniMax and Claude recorded more substantial losses, DeepSeek largest loss -15.10%. The finding does not align with reasoning benchmarks — models that excel at Q&A or code don't necessarily achieve best market performance.
2. More Trades ≠ Better Returns
DeepSeek V3.2, Claude Opus 4.6, Gemini 3.1 Pro Preview each executed >1,000 trades. Grok-4.1 Fast executed ~200 trades. Qwen3.5 Plus, Kimi K2.5, Seed-2.0-Lite recorded 500-800 trades. Quality > quantity. In swing terms: fewer, higher-conviction setups beat hyperactivity. Every extra trade adds spread cost and decision error.
3. Different Models Exhibit Distinct Risk Preferences
Gemini 3.1 Pro Preview and DeepSeek V3.2 used relatively high leverage and suffered larger drawdowns, while Kimi K2.5 managed risk more carefully and achieved strongest return. GPT-5.4 kept exposure low, limiting both gains and volatility. Overall, taking more risk did not necessarily lead to better performance.
| Model Cluster | Return | Leverage / Exposure | Swing Lesson |
|---|---|---|---|
| Qwen / Kimi / Seed — Winners | +~10% / strong positive | Moderate, careful sizing | Risk-adjusted returns win over 6 weeks |
| GPT / GLM — Near Break-Even | ~0% | Low exposure | Capital preservation is a valid strategy in choppy FX |
| DeepSeek / Gemini / Claude — Losers in this window | -15.1% to substantial loss | High leverage, high frequency | Leverage amplifies bad decisions, not edge |
| Grok-4.1 Fast — Efficient | Mid-pack but low trades | ~200 trades only | Conviction filter reduces costs |
Why Qwen Won: Deconstructed
Based on HKU report and agent logs:
- Information synthesis over prediction: Qwen's logs show more web search verification before trade, less reliance on internal memory. It acted like an analyst confirming catalyst, not oracle predicting price.
- Position management: Rather than averaging down losers (common LLM failure mode), Qwen trimmed losing positions early — a risk rule retail traders must hard-code.
- Session awareness: Qwen avoided overtrading during low-volatility Asia lunch, focusing on London/NY overlap — a swing trader equivalent to "only trade when catalyst + volume."
Video Lab — Part 2: Build the Risk-Managed Agent
1. How I Built Blueprint for AI Tradebot — Docker & Python Foundation
Sets up scalable Docker, Flask REST API, and personality layer — foundation for running Qwen-style agents locally.
2. How I Build AI Trading Team in Claude Cowork — No-Code Research Team
Level playing field: institutional teams have analysts, you have a chart. This builds free automated team with FRED, Yahoo Finance connectors — structured data, not hallucinations.
3. Claude AI Agents Are About To Change Crypto Trading — Managed Agents Beta
Anthropic's public beta turning Claude from chatbot into agent infrastructure — tasks, tools, vaults, scheduled runs. Directly relevant to HKU's managed agent concept.
Practical Swing Translation: Your Own Agentic Trader Rules
Steal the winners' behavior without coding full agents:
Rule 1: Fewer, Higher-Conviction Trades
HKU: Grok ~200 trades vs DeepSeek 1,000+. Implement: Max 5 swing positions at once, max 10 new trades per month. Force yourself to rank opportunities 1-10 and only take top 3.
Rule 2: Mandatory Search Verification
Before entry, LLM must fetch 2 independent sources confirming catalyst (e.g., earnings press release + analyst revision). Use prompt:
"Verify today's NVDA decline: Separate temporary sentiment factors from changes to earnings thesis. Compare reaction with historical post-earnings moves (3 years). Identify 5 rebound catalysts and 5 downside catalysts. What evidence would invalidate bullish thesis? Cite sources."
Rule 3: Leverage Cap + Stop Loss = Position Size Math
Formula: Position Size = (Account * Risk%) / (Entry - Stop). Example: $20k account, 0.5% risk = $100, stop $2 away = 50 shares. Never risk more than 0.5% even if thesis is strong. HKU losers ignored this.
Rule 4: Continuous Reassessment
HKU limitation: 6 weeks snapshot only. Market conditions change. Set automated alert: If underlying catalyst changes (guidance cut, FDA rejection, CEO departure), LLM re-evaluates and auto-closes if thesis broken.
What Part 2 Proves About AI-Augmented Swing Trading
Three insights for retail:
- Better reasoning ≠ better trading. Traditional benchmarks insufficient for dynamic markets. Future evaluation must emphasize long-term decision-making in uncertainty.
- Risk control is alpha. Kimi's careful risk management beat high-leverage approaches. In swing, surviving drawdown matters more than picking winner.
- System design > model selection. Qwen vs DeepSeek spread shows model choice matters, but your rules (max exposure, trade frequency cap, verification steps) matter more.
FAQ — Part 2
Is HKU Agentic Trader real money?
Yes, live foreign exchange market data with real spreads and time pressure, $100k identical starting capital per model, identical tools and leverage. Not simulated backtest, though duration is only 6 weeks — snapshot, not long-term investment capability.
Why did Claude lose in HKU but win in Autopilot?
Different environments and oversight. Autopilot Claude had human risk overlay and stock market (vs FX), while HKU was fully autonomous FX with high trade count (1,000+). Shows model performance varies by market, prompt, risk settings — no single model always wins.
Should I run Qwen3 locally via Ollama like in the videos?
Yes for research, paper trading first. TradingAgents repo runs 100% offline on Qwen3 8B on 8GB GPU (RTX 4060). Use it for BUY/SELL/HOLD reports with stop-loss and target, not direct live execution until you add risk caps.
Transition to Part 3: We now know autonomous agents can diverge wildly — +9.9% to -15.1% in same market. But is even the winner's +10% real alpha, or just high-beta momentum exposure? In Part 3, we apply the NBER lens: how to strip factor exposure from LLM picks, why Gemini failed to beat equal-weight, and how to build a factor-neutral swing system.
[Part 2 Complete. Say "Go" or "Proceed" to generate Part 3.]
Part 3: The NBER Factor Trap — Why Claude's +19% and Qwen's +10% Might Be Beta, Not Alpha
PART 3 OF 10 — FACTOR CHECK Part 1: Claude +19.04% vs 12.24% S&P with $27M copy capital. Part 2: Qwen +9.9% vs DeepSeek -15.1% in live FX. Both look like AI intelligence. The 2026 NBER paper "AI Managed Household Portfolios" says: Not so fast. Strip large-cap, high momentum, high beta, low book-to-market exposure — and abnormal returns vanish.
This is how to tell if your LLM is smart — or just riding momentum.
What NBER Actually Did — Prospective, Not Backtest
Unlike most viral "ChatGPT beats market 500%" threads, NBER researchers collected prospective daily recommendations — they asked LLMs each day to build portfolios going forward, then tracked actual next-day returns. No lookahead, no training data leakage. Sample: ChatGPT-5, Grok 4.1 Fast, Claude Sonnet 4.5, Gemini 2.5 Flash, DeepSeek, etc. Two portfolio types: passively managed buy-and-hold and actively managed.
Finding 1: Investors turning to AI chatbots for stock picks may end up with portfolios heavily concentrated in Big Tech and semiconductor stocks — without gaining meaningful edge over market.
Finding 2: AI portfolios tilted heavily to:
- Large market capitalization (Mega-cap bias)
- Strong recent momentum (last 3-12 month winners)
- High beta (moves more than market)
- Low book-to-market (growth over value)
- Heavy media coverage (availability bias)
Finding 3: Once returns were adjusted for Fama-French 5 factors + momentum (market, size, value, profitability, investment, momentum), buy-and-hold and actively managed AI portfolios did not show statistically significant abnormal returns.
Why This Matters for Your Swing Trading
Imagine this swing scenario, common in 2025-2026:
LLM suggests long NVDA after -7% post-earnings dip. Thesis: Temporary sentiment, not thesis change, historical bounce, 3 catalysts. NVDA bounces +12% in 2 weeks. You profit.
Is that AI alpha? Or did you just buy high-beta momentum that bounces harder when market bounces? NBER says you must ask:
- Would an equal-weight Big Tech + semiconductors basket have done same? If yes, no alpha.
- Would a simple 50-day momentum screen have picked same names? If yes, no alpha.
- Did AI add timing beyond factor? Did it pick entry day better than random within 5-day window?
Real Example: Claude Portfolio vs Factor
Claude Portfolio launched March with $50k, +19.04% by Aug 4 vs 12.24% S&P. On surface, +6.8% alpha. But if Claude held 80% in 7 mega-cap tech names during a period where Nasdaq-100 was +15%, and beta was 1.3, then expected return = 12.24% * 1.3 = ~15.9%. Remaining alpha ~3.1%, before fees/slippage. Still positive, but far from headline 19%.
This doesn't mean Claude failed — it means retail must benchmark correctly.
How to Benchmark Correctly — The Retail Checklist
- Compare to sector: If AI picks NVDA, compare to SMH (semiconductor ETF), not just SPY
- Compare to equal-weight: Equal-weight version of AI's top 10 holdings
- Calculate beta: If your AI portfolio beta = 1.4, multiply market return by 1.4 for expected
- Check momentum overlap: What % of AI picks were top 20% momentum last 6 months?
- Check concentration: Are 5 of 10 picks Big Tech + semis? If yes, heavy concentration risk
- Timing test: Would buying same names 3 days later have similar return? If yes, timing alpha weak
- Drawdown-adjusted: Sharpe ratio, not raw return. HKU: high risk ≠ better returns
| Metric | What Naive Retail Does | What Pro Does (NBER-Style) |
|---|---|---|
| Benchmark | SPY only | SPY + Sector ETF + Equal-Weight + Factor-adjusted |
| Return | +19.04% headline | +19.04% minus beta-adjusted market, minus fees/slippage |
| Concentration | Ignores | Measures % in top 7 mega-caps, Herfindahl index |
| Momentum bias | Ignores | Measures overlap with 12-2 momentum factor |
| Trade frequency | More trades = more skill | HKU: 200 trades (Grok) more efficient than 1,000+ (DeepSeek) |
Building a Factor-Neutral LLM System
Step 1: Force Sector Diversification in Prompt
Instead of "Pick 10 best stocks," use:
"Build 10-stock portfolio: Max 2 per sector (Tech, Semis, Healthcare, Energy, Industrials, Financials, Consumer), max 30% mega-cap (> $500B), must include 2 mid-cap ($5-50B), must include 1 low-momentum turnaround. Explain factor exposures."
Step 2: Ask LLM to Self-Report Factors
Add: "For each pick, report: market cap, 6-month momentum percentile, beta vs SPY, book-to-market percentile, and whether it's in top 50 most-covered news names. Then calculate portfolio averages."
This forces transparency — the model must admit if it's just picking Nvidia again.
Step 3: Use Investment Committee Model
TradingAgents architecture (53k stars) uses 7 agents: 4 analysts (fundamental, sentiment, news, technical), 2 debaters (bull vs bear), trader, risk team, portfolio manager. Retail version:
- ChatGPT / DeepSeek: Fundamental thesis + code
- Claude: Adversarial risk review — "Find every reason this fails"
- Grok / Qwen: News sentiment + rapid catalyst detection
- Risk manager: Position size, max loss, sector caps, factor exposure
Video Lab — Part 3: From Single Pick to Factor-Aware System
1. Someone Open-Sourced a Hedge Fund — Running Locally on Qwen3
Why Qwen3 8B on Ollama on 8GB GPU matters for factor-neutral testing — no API keys, no cost, reproducible. Full BUY/SELL/HOLD report with stop-loss and 3-6 month target.
2. How I Run Multiple Claude AI Trading Agents on Autopilot — Competition
Running multi-agent strategy competition — exactly how to test factor exposure: run agents with different sector constraints and compare risk-adjusted returns, not raw returns.
3. How to Trade Using AI: Step-by-Step Strategy & Risk Management
CFA/Columbia workflow: sentiment analysis to quantify animal spirits, deep learning for intrinsic value using Nissim and Penman, and verification checklist against hallucinations — essential for NBER-style rigor.
Original Analysis: What Happens Next
What happened: NBER exposed that LLM portfolios' impressive headline returns are largely explained by factor tilts — large, high-momentum, high-beta growth.
Why it matters: Retail investors who chase AI picks without factor adjustment will overpay for beta and suffer when momentum regime flips. March 2025 - May 2026 was momentum regime; next regime may punish same tilts.
Who benefits if fixed: Traders who force sector diversification, cap mega-cap, require mid-cap inclusion, and benchmark against equal-weight and sector. They will discover true AI edge is in timing and risk veto, not stock universe.
Risks: Over-constraining kills edge; under-constraining is closet momentum fund with LLM fees. Also, FINRA warns AI can generate false information — factor reports must be verified with Bloomberg/CapIQ, not just LLM.
What's next: Next generation retail tools will auto-calculate factor exposures and show "Alpha vs Beta" dashboard — like Claude Finance Agents already starting to do. Edge shifts from "what to buy" to "why it beats factor benchmark."
FAQ — Part 3
Does NBER prove AI portfolios have no alpha?
No. It proves headline returns are largely explained by factors, and abnormal returns after adjustment were not statistically significant in their sample/methodology. It doesn't prove zero alpha in all regimes — it proves you must adjust for factors to know if alpha exists.
How do I quickly check factor exposure without Bloomberg?
Ask LLM to output for each holding: market cap, 6m momentum percentile, beta (Yahoo Finance), and sector. Average them. If 80% are >$500B, top 20% momentum, beta >1.3, tech/semis, you are high-beta momentum fund.
Is equal-weight better than AI?
Often equal-weight of AI's own top picks beats cap-weighted AI portfolio on Sharpe, because it reduces mega-cap concentration. Test both: ask AI for 10 picks, then backtest equal-weight vs AI-weighted vs SPY vs sector ETF.
Transition to Part 4: We now know how to strip beta from AI returns. In Part 4, we build the system that survives that test — the open-source TradingAgents hedge fund with 7 agents, bull vs bear debate loop, trader, risk team, and PM. You'll get the exact prompts to run it locally on Qwen3 via Ollama on your RTX 4060, with crash recovery and persistent memory, and how to enforce factor caps inside the agent prompts.
[Part 3 Complete. Say "Go" or "Proceed" to generate Part 4.]
Part 4: TradingAgents — Inside the 53,000-Star Open-Source Hedge Fund Retail Investors Run on Qwen3
PART 4 OF 10 — BUILD LAB From theory to terminal. Part 2 showed autonomous agents diverging +9.9% to -15.1%. Part 3 showed factor exposure explains most headline alpha. Part 4 gives you the architecture that survives both tests: TradingAgents — 7 specialized LLM agents (4 analysts, bull/bear debaters, trader, risk team, portfolio manager), 53k+ GitHub stars, Apache 2.0, built on LangGraph, with v0.2.4 crash recovery and persistent memory.
You'll learn how retail runs it 100% offline on Ollama with Qwen3 8B on an 8GB RTX 4060 — full BUY/SELL/HOLD report with stop-loss, position sizing, 3-6 month target — and how to inject factor caps from Part 3 directly into agent prompts.
The Architecture: How Real Hedge Funds Work, Mirrored in LLMs
Real hedge funds aren't one genius — they're a team with adversarial debate. TradingAgents (TauricResearch/TradingAgents, paper arXiv:2412.20138, UCLA) mirrors it:
Fundamentals (10-K/10-Q, Nissim & Penman profitability), Sentiment (NLP on news/social), News (catalyst detection), Technicals (RSI, MACD, Bollinger, volume, ATR, short interest). Each produces structured report.
Two researcher agents debate each other with evidence. This is the key design choice — prevents echo chamber. Bull argues upside catalysts, bear finds every reason it fails. Debate transcript becomes input to trader.
| Agent | Role | What It Prevents |
|---|---|---|
| Fundamental Analyst | Estimates intrinsic value, profitability forecast | Buying hype without cash flow |
| Sentiment Analyst | Quantifies animal spirits from news/social | Missing panic/euphoria shift |
| News Analyst | Detects catalysts, guidance changes | Trading on stale thesis |
| Technical Analyst | High-probability setups, support/resistance | Bad entry timing |
| Bull Researcher | Best upside case with data | Missing opportunity |
| Bear Researcher | Best downside case with data — veto power | Blowup from ignored risk |
| Trader + Risk + PM | Final BUY/SELL/HOLD with sizing, stop, target | Over-leverage, no exit plan |
How Retail Runs It Locally — No API Keys, No Cost
Stack Used (from video lab)
TradingAgents repo: github.com/TauricResearch/TradingAgents Ollama — Local LLM runner Qwen3 8B — 8GB GPU friendly, Alibaba's strongest in HKU +9.9% LangGraph — Under the hood orchestration Miniconda — Python 3.13 env RTX 4060 8GB — Runs full report ~8-12 minutes per ticker
Setup Steps
git clone https://github.com/TauricResearch/TradingAgents.git cd TradingAgents conda create -n tradingagents python=3.13 conda activate tradingagents pip install -r requirements.txt ollama pull qwen3:8b python -m tradingagents.cli --ticker NVDA --date 2026-05-13 --analysts 4 --llm qwen3:8b
What you get: Fundamentals → Sentiment → News → Technicals → Bull argument → Bear argument → Trader decision → Risk team adjustment → Portfolio manager final: BUY/SELL/HOLD, stop-loss, position size %, 3-6 month price target.
v0.2.4 upgrade: Structured-output decision agents (JSON, not prose), Docker support, crash recovery and persistent memory — so a 2-hour NVDA run doesn't lose progress if laptop sleeps.
Injecting Part 3 Factor Caps Into TradingAgents Prompts
Default TradingAgents can still pick 80% Big Tech. Fix it by editing default_config.py and analyst prompts:
SYSTEM_PROMPT_ADDITION = """ CONSTRAINTS — Factor-neutral: - Max 2 stocks per sector (Tech, Semis, Healthcare, Energy, Industrials, Financials, Consumer, Staples) - Max 30% mega-cap > $500B, must include 2 mid-cap $5-50B, must include 1 low-momentum turnaround (6m momentum < 20th percentile) - For each pick, report: market cap, 6m momentum percentile, beta vs SPY, book-to-market percentile, news coverage rank - If portfolio beta >1.3, reduce position size by 30% and add hedge - Bear researcher has veto if concentration >40% in one sector """
This implements NBER lesson directly: force diversification and transparency.
Video Lab — Part 4: Watch It Run
1. Someone Open-Sourced a Hedge Fund — 53k Stars Explained
Deep dive on what's inside, why bull vs bear debate matters, spec sheet, 4 analysts, installation, verdict. Repo link, paper link, chapters.
2. I Ran It Locally on Qwen3 with Ollama — RTX 4060 Demo
NVDA report live: ticker, date, analysts, LLM choice, why Qwen3 8B for 8GB GPU, reports breakdown fundamentals → portfolio manager, final decision target $268, 3-6 months, honest limitations.
3. How I Build AI Trading Team in Claude Cowork — No-Code Version
For non-coders: build free automated AI trading team using Claude Cowork data connectors (FRED, Yahoo Finance), setup agents in cowork, structured data not hallucinations. Same team concept, no Python.
From Research to Swing Execution — The Workflow
TradingAgents output is research, not execution. Translate to swing:
- Screen: Deterministic screener — RSI <35 with volume 2x avg, or earnings surprise >5%.
- Run TradingAgents: Feed ticker + date. Get BUY/SELL/HOLD, stop, target.
- Multi-model check: Run same ticker with different LLMs (Qwen3 vs Claude vs DeepSeek) via
--llmflag. If 2 of 3 say BUY with similar stop, conviction higher. - Risk math: Use 0.5% risk rule: Position = (Account * 0.005) / (Entry - Stop). If TradingAgents suggests stop $4 away on $20k account, position = $100 / $4 = 25 shares.
- Human approval: Check bear researcher's veto reasons — FDA date, earnings in 2 days, CEO selling. If bear reason is temporal, wait.
- Journal: Log bull/bear debate transcript, not just final decision. Review weekly: Did bear veto save you?
Risks & Limitations
- Not financial advice: TradingAgents README explicitly says research framework, exchange piece simulated, no live broker wired by default.
- Lookahead risk: Even with live data, LLM training cutoffs may include info about ticker. Use date param to force point-in-time.
- Compute cost: Qwen3 8B on RTX 4060 ~8-12 min per ticker, 4 analysts parallel. 20 tickers = ~3 hours. Batch overnight.
- Overfitting debate: Bull/bear debate can produce plausible but false narratives. Require citations — force analysts to quote 10-K section, not summarize.
FAQ — Part 4
Do I need to code to use TradingAgents?
No for no-code path: use Claude Cowork version (PB75sDtkRqs video) — build team with connectors, no Python. Yes for full TradingAgents — clone repo, conda, Ollama, CLI. Both produce same 7-agent logic.
Why Qwen3 over GPT-4 or Claude?
HKU live FX: Qwen +9.9% best, Claude/GPT near break-even to loss in that window. Qwen3 8B also runs locally on 8GB GPU, no API cost, privacy, reproducible. For factor-neutral testing, local = controllable.
Can TradingAgents place live trades on Alpaca/Blofin?
Repo simulates fills; you must wire brokerage API yourself. Recommended: start with paper trading, log simulated fills to journal/portfolio.json, then add Alpaca API with max exposure 70% and 0.5% risk cap. Never give agent unlimited API keys.
Transition to Part 5: We have the team — 7 agents debating with risk veto. In Part 5, we put them to work on the exact workflow retail profits came from: Screen → Catalyst → Multi-model debate → Technical → Fundamental → Risk → Human Approval → Monitoring → Journal. You'll get copy-paste prompts for catalyst-driven swing trading — earnings reversals, guidance changes, sector rotation — with entry/exit templates and the 0.5% risk calculator.
[Part 4 Complete. Say "Go" or "Proceed" to generate Part 5.]
Advanced Prompt Injection for Factor Caps and Stop-Loss Discipline
Retail failure mode: agent says BUY but no stop. Force it. Edit tradingagents/prompts/trader_prompt.py:
TRADER_PROMPT = """
You are a swing trader with strict discipline.
You MUST output JSON: {"action": "BUY/SELL/HOLD", "entry": float, "stop": float, "target": float, "position_pct": float, "reasoning": str, "invalidation": str}
Rules:
- Stop must be 1.5*ATR or recent swing low, whichever is tighter
- Position_pct = min( (0.005*account)/(entry-stop), 0.15 ) — max 15% per trade
- Invalidation: What news would make you exit immediately?
- Factor check: If sector concentration >40% after this trade, reduce position_pct by 50%
- No trade if bear researcher veto contains time-sensitive risk (earnings in <3 days, FDA date <5 days)
"""
This JSON output is v0.2.4 structured-output feature — it prevents prose-only hallucinations. You can pipe JSON directly to your journal or TradingView webhook.
Combine with Part 3 checklist: After each TradingAgents run, run second LLM pass: "Given this portfolio, calculate sector %, mega-cap %, momentum %, beta. Flag if any > threshold." That second pass is your risk team — exactly HKU's missing piece for losers who used high leverage.
Part 5: The Workflow That Actually Profits — From Screen to Journal with AI Catalyst Trading
PART 5 OF 10 — WORKFLOW Parts 1-4 gave you proof and architecture. Claude +19.04% vs 12.24% S&P with $27M copy, Qwen +9.9% vs DeepSeek -15.1% in live FX, factor trap exposed, TradingAgents 7-agent system. Part 5 answers: How do retail investors actually turn that into daily swing profits?
The answer from profitable retail in 2026: Information compression > stock picking. Not "pick me a winner" but Screen → Catalyst → Multi-model debate → Technical → Fundamental → Risk → Human Approval → Monitoring → Journal. With hard risk caps.
The 9-Step AI-Augmented Swing System
Step 1: Screen — Deterministic, Not LLM
Don't ask LLM to screen — use deterministic screener. Examples that worked in 2026 case studies:
- Post-earnings reversal: Earnings surprise >+5% but stock -5% day 1, volume 2x avg, RSI <45
- Guidance raise + price flat: Guidance up >8%, price change last 5 days <2%
- Sector rotation: Sector ETF +3% week, stock in sector flat, short interest >12%
Output: 5-10 candidates, not 100. LLMs are bad at scanning 5000 tickers, good at deep dive on 5.
Step 2: Catalyst — What Changed?
This is LLM superpower. Prompt:
For ticker {TICKER}, date {DATE}, list:
- What changed in last 7 days? (earnings, guidance, contract, insider, regulatory, macro)
- Separate temporary sentiment vs permanent thesis change
- Compare reaction to historical reactions for same catalyst type (3-year lookback)
- 5 rebound catalysts and 5 downside catalysts with probability %
- What would invalidate bullish thesis? Cite 2 primary sources.
Step 3: Multi-Model Debate — Investment Committee
Run same prompt through Qwen3, Claude, GPT-5, DeepSeek. HKU showed models have distinct risk preferences. You want disagreement. If 3 of 4 agree BUY with similar stop, conviction high. If 2-2 split, skip.
TradingAgents already does bull vs bear debate. For swing, add third agent: Risk Veto — "You are bear researcher with veto if earnings in <3 days, FDA <5 days, concentration >40% sector, beta >1.4 without hedge."
Fundamental thesis + code for backtest, estimates intrinsic value using Nissim & Penman
Adversarial risk review, factor exposure report (market cap, momentum, beta, coverage)
Step 4: Technical — Structured Data In, Not Vague
Don't say "What do you think of NVDA chart?" Feed structured:
Price: $42.80, 5d +8.4%, 20d +18.7%, 50MA $39.20, RSI 68, Volume 2.4x avg, ATR $1.62, Short 14%, Earnings 23d, Beta 1.35
Then ask: Trend → Momentum → Volatility → Support/Resistance → Catalyst confirmation → Risk/Reward. This prevents hallucination.
Step 5: Fundamental — Quick Verification
Use Claude Finance Agents (video in Part 1). Prompt: "Pull 10-Q last quarter revenue growth, gross margin trend, guidance vs consensus. Is thesis supported by numbers or just story?"
Step 6: Risk — The 0.5% Rule That Saved Qwen
HKU winners managed risk carefully, losers used high leverage. Hard rule:
- Max 0.5% account risk per trade
- Max 15% position per trade
- Max 70% total exposure
- Max 3x leverage
Formula: Shares = (Account * 0.005) / (Entry - Stop)
Account Size ($): Entry Price ($): Stop Price ($):
Step 7: Human Approval — The Final Veto
AI proposes, human disposes. Checklist before click:
- Bear researcher veto? If yes, skip or reduce size 50%
- Earnings/FDA in next 3 days? If yes, skip
- Factor exposure: Is this 5th tech/semis trade? If yes, require mid-cap non-tech instead
- Journal: Have I lost 2 in a row with same setup? If yes, pause 24h (HKU: more trades ≠ better)
Step 8: Monitoring — Automated Thesis Check
Set LLM to re-evaluate every day at close: "Has invalidation condition from Step 2 occurred? Guidance cut? CEO selling? Sector breakdown?" If yes, auto-close. This is agentic trading — research is first use case, not just entry.
Step 9: Journal — Where Edge Compounds
Log: Ticker, date, catalyst, bull/bear debate transcript, factor report, entry/stop/target, risk %, outcome, what bear got right. Review weekly: Did bear veto save you? Did factor cap help? This is how retail builds edge that compounds.
Catalyst Templates That Worked in 2026
| Catalyst Type | LLM Prompt Hook | Example Entry | Exit |
|---|---|---|---|
| Earnings Reversal | Separate sentiment vs thesis, compare to 3-year post-earnings bounce rate | Day 2 after -5% on +5% surprise, RSI <40, volume 2x | 50% at 50MA, rest at bear invalidation or 20d high |
| Guidance Raise + Flat Price | Find guidance raise >8% with price <2% 5d, identify why market ignored | Break of Day 1 high post-guidance with volume | Stop under guidance day low, target recent swing high |
| Sector Rotation | Sector ETF +3% week, stock flat, short interest >12%, catalyst? | First close above 20d MA with sector confirmation | Sector ETF breaks 10d low |
Video Lab — Part 5: Watch the Workflow Built
1. ChatGPT Just Changed the Stock Market Forever — SIGNAL Framework
Level 1 setup bot, proof of results, trailing stop, Level 2 smart money signals, copy trading politicians, trading from phone. Full SIGNAL framework.
2. Claude's New Trading Agent Is Insane — Smart Money + Copy Bot
Setup, trade with smart money, building copy trading bot from any source — Level 2 and Level 3 options trading wheel strategy.
3. I Built a Trading Bot with ChatGPT — $2000 Live 24h
Alpaca API + FinRL + Vercel deployment — context memory advantage for MVP building.
Original Analysis: Why This Workflow Beats Stock Picking
What happened: Retail who asked "What stock should I buy?" got concentrated Big Tech momentum portfolios (NBER) and overtraded 1,000+ times like DeepSeek in HKU, losing -15.1%.
Why it matters: Retail who implemented 9-step workflow with 0.5% risk, factor caps, bear veto, and journaling replicated Qwen's careful risk management and achieved positive risk-adjusted returns even in choppy FX.
Who benefits: Traders who accept AI is research department, not oracle. Information compression (10-K 7k words → 5 bullets) + discipline (max 5 positions, 10 trades/month) = edge.
Risks: Over-automation without human veto (FINRA: AI can generate false info), ignoring factor exposure, leverage creep. Also, Grok 200 trades efficient vs 1,000+ inefficient — more automation ≠ better.
What's next: Claude Managed Agents beta + Coinbase for Agents + Blofin API = retail can run 24/7 assistant that controls desktop, analyzes journals, executes via API with vaulted credentials. Edge moves to who builds best system around agents, not best indicator.
FAQ — Part 5
How many trades per month is optimal?
HKU: Grok ~200 trades over 6 weeks (~33/month) more efficient than 1,000+ (~166/month). For swing, 5-10 trades/month max with strict catalyst filter. Quality > quantity. Every trade adds spread and error.
What's the best prompt for bear researcher veto?
"You are bear researcher with veto power. Find every reason this swing trade fails: earnings in <3 days, FDA <5 days, insider selling, sector breakdown, factor concentration >40%, beta >1.4 without hedge. If any true, output VETO: [reason]. If not, output NO VETO with 2 residual risks."
Can I run this without coding?
Yes: Claude Cowork no-code team (video PB75sDtkRqs) with FRED + Yahoo Finance connectors, plus n8n workflow Part 2 (build AI-powered trading bot). For full 7-agent system, need TradingAgents repo + Ollama.
Transition to Part 6: We have the system — 9 steps, risk calculator, bear veto. In Part 6, we apply it to the most profitable swing setup of 2026: post-earnings reversal. You'll get 3 live examples (NVDA -7% case, guidance raise flat price, sector rotation squeeze), with exact TradingAgents reports, entry/stop/target math, and how Claude Finance Agents pull 10-Q in under an hour.
[Part 5 Complete. Say "Go" or "Proceed" to generate Part 6.]
Part 6: Post-Earnings Reversal — The Most Profitable Swing Setup of 2026 with AI Catalyst Verification
PART 6 OF 10 — CATALYST LAB Parts 1-5: Real-money proof (Claude +19.04% vs 12.24% S&P, $27M copy), live agent divergence (Qwen +9.9% to DeepSeek -15.1%), factor trap (NBER), TradingAgents 7-agent system, 9-step workflow with 0.5% risk calculator. Part 6: We apply it to the setup that produced highest hit rate in 2026 retail journals — post-earnings reversal — where price drops -5% on +5% earnings surprise, but thesis unchanged.
You'll get 3 live examples, exact prompts, TradingAgents outputs, and how Claude Finance Agents pull 10-Q in under an hour with no coding.
Why Post-Earnings Reversal Works for LLMs
Earnings are noisy. Market reacts to headline EPS vs whisper, not underlying thesis. LLMs excel at separating temporary sentiment from permanent thesis change because they can read full 10-Q (7,000+ words), transcript, guidance, and compare to 3-year historical post-earnings bounce pattern — something human can't do in 10 minutes before close.
Three conditions that define edge:
- Surprise >+5% EPS vs consensus, but price -5% day 1 — market disappointed by guidance nuance, not core
- Volume 2x average, RSI <40 — capitulation, not drift
- Thesis intact: Revenue growth, gross margin, guidance vs consensus still up
Template: Claude Finance Agents Pull 10-Q in Under an Hour
From video "Don't Analyse Stocks Without Claude's New Finance Agents" — no coding, no spreadsheets, no data subscriptions:
Step 1: Install Claude Marketplace → Finance Plugins (Fundamentals, Sentiment, News, Technicals)
Step 2: Command: "/analyze {TICKER} earnings {DATE}"
Step 3: Claude auto-pulls 10-Q, transcript, guidance, analyst revisions
Step 4: Output: Revenue growth, margin trend, guidance vs consensus, 5 catalysts, 5 risks, factor report
Time: 45-60 minutes vs 3-4 hours manual
Live Example 1: NVDA -7% After +6% Surprise — The Classic
Ticker: NVDA — Date: Simulated May 13, 2026
Screen: EPS +6% vs consensus, revenue +12% YoY, stock -7% day 1, volume 2.4x, RSI 38, ATR $1.62
TradingAgents Report (Qwen3 8B local):
- Fundamental: Revenue growth intact, data center +18%, gross margin 74.5% +80bps YoY
- Sentiment: News negative on CFO OpEx timing comment, but 8 of 10 analysts raised target
- News: No regulatory change, no insider selling spike
- Technical: Support at $41.20 (recent swing low), resistance $45.80
- Bull: Historical post-earnings bounce 68% of time when surprise >5% and price -5%, avg bounce +9% in 10 days
- Bear: OpEx timing could compress Q3 margin, beta 1.35 high, sector concentration risk if already 3 semis
- Trader: BUY $42.80, Stop $40.80 (swing low -1*ATR), Target $46.50 (recent high), Position 0.5% risk
Risk Math: $20k account, 0.5% = $100, Entry $42.80 - Stop $40.80 = $2 → 50 shares → $2,140 position (10.7% of account) — under 15% cap OK.
Outcome template: Day 2 close above Day 1 high with volume confirmation = add 25% if bear veto cleared. Exit 50% at 50MA, rest at bear invalidation or 20-day high.
Live Example 2: Guidance Raise + Flat Price — The Ignored Catalyst
Ticker: Mid-cap Industrials — Guidance +9% vs consensus, Price +0.8% 5 days
Screen: Guidance raise 9% vs consensus, price change 5d 0.8%, volume flat, short interest 13%
LLM Prompt: "Find guidance raise >8% with price <2% 5 days. Why did market ignore? Compare to historical guidance raise reactions for this ticker (3 years). Identify if low coverage or sector rotation distraction."
TradingAgents Bear Veto Check: No earnings in <3 days, no FDA, sector concentration 20% only, beta 1.1 — NO VETO, 2 residual risks: Industrials sector ETF lagging, insider sale $1.2M last week but not cluster.
Entry: Break of guidance day high $54.20 with volume 1.8x. Stop under guidance day low $51.90. Target recent swing high $58.00.
Risk Math: $20k *0.5%=$100 / ($54.20-$51.90=$2.30) = 43 shares → $2,330 (11.6%) — OK.
Live Example 3: Sector Rotation Squeeze
Ticker: Short Interest >12% in Sector ETF +3% Week, Stock Flat
Screen: XLI Industrials +3.2% week, stock in XLI flat -0.2%, short interest 14%, RSI 52
LLM Prompt: "Sector ETF +3% week, stock flat, short >12%. Is there catalyst? Earnings? Contract? If no news, is it short squeeze candidate? Compare short interest vs 1-year avg, days to cover, borrow rate. List 3 squeeze catalysts."
TradingAgents: No fundamental change, but technical: 20-day MA curling up, first close above 20d with volume, bear argument weak (only "sector overbought"). Trader: BUY $38.50, Stop $36.80 (below 20d), Target $41.00 (short interest trigger).
Factor Check (from Part 3): Is this 5th industrials trade? If yes, require non-industrials next. Mega-cap %? Must include mid-cap. This prevents Big Tech concentration NBER warned about.
Video Lab — Part 6: Watch Catalyst System Built
1. Don't Analyse Stocks Without Claude's Finance Agents — Full Install
Step-by-step install, which plugins to install/skip, how to research stock in under an hour with no coding. Chartered Accountant ex-private equity workflow.
2. How to Trade Using AI: Step-by-Step Strategy & Risk Management
CFA Institute + Columbia workflow: NLP sentiment to quantify animal spirits, deep learning intrinsic value, seasonality + technical confirmation, Black Swan checklist against hallucination.
3. Agentic Trading: A New Way To Trade — Research First Use Case
Nansen CEO Alex Svanevik: why research is first use case for AI agents, trust through transparency, managing hallucinations, who benefits most.
Original Analysis: What Happened, Why It Matters, What's Next
What happened: Retail who chased headline "AI picks NVDA" without catalyst verification got whipsawed when guidance nuance mattered. Those who used LLM to separate sentiment vs thesis and waited Day 2 confirmation captured +8-12% bounce with defined risk.
Why it matters: Post-earnings reversal is highest hit-rate swing setup in 2026 journals because it exploits behavioral bias — market overreacts to tone, LLM reads numbers. But only works with 0.5% risk rule and bear veto (earnings in <3 days skip).
Who benefits: Traders who implement 3-day rule: Day 1 earnings, Day 2 verification with TradingAgents, Day 3 entry only if bear veto cleared. Prevents FOMO entry into -7% falling knife.
Risks: Earnings drift can continue if thesis broken (revenue miss, not just guidance tone). Must require fundamental intact: revenue growth + margin trend. Also, HKU: more trades ≠ better — limit to 5 post-earnings setups per month max.
What's next: Claude Cowork Dispatch + scheduled runs will auto-run "/analyze {TICKER} earnings" at 7am next day, push report to Discord/email, with factor exposure calculated. Edge moves to who automates verification, not who reads fastest.
FAQ — Part 6
What if TradingAgents says HOLD but I see reversal pattern?
Trust HOLD if bear veto contains time-sensitive risk (earnings <3d, FDA <5d). If HOLD due to factor concentration (5th tech trade), switch to non-tech candidate. HOLD is valid signal — HKU: near break-even agents preserved capital better than high-leverage losers.
How to avoid buying falling knife on -7% day?
Never Day 1. Use 3-day rule: Day1 earnings, Day2 LLM verification (10-Q + transcript + factor report), Day3 entry only if close above Day1 high with volume 1.5x. Stop under Day1 low. This filters 60% of false reversals.
Can I automate this with n8n?
Yes — Part 2 video "Build AI-Powered Trading Bot Part 2 n8n" shows real workflow chaining stock data + ChatGPT agents. Connect earnings calendar → screener → TradingAgents CLI → Discord alert with BUY/SELL/HOLD + risk math.
Transition to Part 7: We've mastered single catalyst — post-earnings reversal. In Part 7, we scale to technical + structured data feeding: how to feed RSI, ATR, volume, short interest, and 10-Q numbers without hallucination, using Claude Finance Agents and TradingAgents technical analyst. You'll get the structured data prompt template that prevents LLM from inventing prices.
[Part 6 Complete. Say "Go" or "Proceed" to generate Part 7.]
Part 7: Technical Analysis + Structured Data — How to Feed RSI, ATR, Volume, Short Interest Without Hallucination
PART 7 OF 10 — STRUCTURED DATA Parts 1-6: Real-money proof (Claude +19.04% vs 12.24% S&P), live FX divergence (Qwen +9.9% to DeepSeek -15.1%), factor trap, TradingAgents 7-agent build, 9-step workflow, post-earnings reversal with 3 live trades. Part 7: The failure mode that killed DeepSeek and Gemini in HKU — hallucinating prices and ignoring structured risk. Solution: feed LLMs structured JSON, not vague chart talk.
If you say "What do you think of NVDA chart?" you get hallucination. If you feed Price $42.80, RSI 68, ATR $1.62, Vol 2.4x, Short 14%, Beta 1.35 and force JSON output, you get tradable levels.
Why LLMs Hallucinate Technicals — And How HKU Losers Did It
HKU report: Gemini 3.1 Pro Preview and DeepSeek V3.2 used relatively high leverage and suffered larger drawdowns. Why? Logs show they acted on vague sentiment without verifying ATR, stop distance, or position size. They averaged down losers because no structured stop was forced.
Common retail prompt failure:
"Analyze NVDA chart, should I buy?" — Model invents price $38.20 when actual $42.80, invents RSI 45 when actual 68, suggests no stop.
Fix: Structured data block + forced JSON schema from TradingAgents v0.2.4.
The Structured Data Template — Copy-Paste
STRUCTURED_MARKET_DATA = {
"ticker": "NVDA",
"date": "2026-05-13",
"price": 42.80,
"change_5d_pct": 8.4,
"change_20d_pct": 18.7,
"ma_50": 39.20,
"ma_200": 36.10,
"rsi_14": 68,
"atr_14": 1.62,
"volume_multiple": 2.4,
"short_interest_pct": 14.0,
"days_to_cover": 3.2,
"beta_spy": 1.35,
"earnings_days_away": 23,
"sector": "Semiconductors",
"sector_etf_5d_pct": 2.8
}
PROMPT = """
You are technical analyst in TradingAgents. Use ONLY data in STRUCTURED_MARKET_DATA. Do not invent prices.
Tasks:
1. Trend: Price vs 50MA and 200MA
2. Momentum: RSI interpretation, but note RSI 68 is not overbought if volume 2.4x
3. Volatility: ATR-based stop = 1.5*ATR
4. Support/Resistance: Recent swing low $41.20, high $45.80
5. Risk/Reward: Calculate with structured stop
Output JSON ONLY:
{
"trend": "bullish/neutral/bearish",
"setup": "post-earnings reversal / momentum breakout / sector rotation / none",
"entry": float,
"stop": float,
"target": float,
"rr_ratio": float,
"invalidation": "What breaks thesis?",
"confidence": "low/medium/high"
}
"""
entry, stop, target, rr_ratio, invalidation prevents prose hallucination. TradingAgents v0.2.4 added structured-output decision agents exactly for this — retail should copy it. No stop = no trade.
Example: Feeding RSI, ATR, Volume Correctly
| Indicator | Raw Value | What LLM Should Do | What Hallucination Looks Like |
|---|---|---|---|
| RSI 68 | 68, not 45 | Note strong momentum but not extreme, check volume 2.4x confirms | "RSI oversold at 30" — invented |
| ATR $1.62 | $1.62 | Stop = 1.5*ATR = $2.43 below entry → $40.37 | Stop $2 random |
| Volume 2.4x | 2.4x avg | Capitulation or breakout confirmation | "Low volume" — false |
| Short 14% | 14%, days to cover 3.2 | Squeeze potential if sector +3% | Ignores short |
| Beta 1.35 | High beta | Expect 1.35x market move, adjust position -30% if portfolio beta >1.3 | Treats as low beta |
Resulting JSON from Qwen3:
{
"trend": "bullish",
"setup": "post-earnings reversal",
"entry": 42.80,
"stop": 40.37,
"target": 46.50,
"rr_ratio": 1.52,
"invalidation": "Close below $41.20 swing low or sector ETF breaks 10d low",
"confidence": "medium"
}
Position Size from ATR Stop
Use 0.5% rule from Part 5: $20k account *0.005 = $100 risk. Entry $42.80 - Stop $40.37 = $2.43 → 41 shares → $1,754 position (8.7% of account) — under 15% cap, passes factor check.
Video Lab — Part 7: Building Structured Pipelines
1. How I Built Blueprint for AI Tradebot — Docker Foundation
Episode 1 architecture: Dockerized Python, Flask REST API, pixel art bots — foundation for structured data feeding via REST, not chat.
2. Build AI-Powered Trading Bot Part 2 — n8n Real Workflow
Actual n8n workflow: configure multiple workflows, connect trading APIs, chain real-time stock data analysis with ChatGPT agents — structured data chain.
3. How I Build AI Trading Team in Claude Cowork — Data Connectors
FRED + Yahoo Finance connectors to stop hallucinations — structured data, not hyped headlines. Accuracy when real money on line.
Integrating with Claude Finance Agents
From Part 1 video: Claude Marketplace → Finance Plugins → Command /analyze {TICKER} earnings {DATE} pulls 10-Q in under hour. Combine with structured technical block:
Step 1: /analyze NVDA earnings 2026-05-13 → fundamentals Step 2: Feed STRUCTURED_MARKET_DATA JSON → technicals Step 3: Prompt: "Combine fundamental from Step1 + technical from Step2, output JSON with entry/stop/target/invalidation/confidence/factor report" Step 4: Risk team prompt: "Given portfolio sector %, beta, mega-cap %, does this trade breach factor caps from Part3? If yes, reduce position 50%."
This 4-step chain is exactly what TradingAgents does: 4 analysts parallel → bull/bear debate → trader → risk → PM. Retail can replicate with Claude Projects.
FAQ — Part 7
Why not just paste TradingView screenshot to LLM?
Vision models hallucinate levels. They read $38.20 as $42.80, misread RSI scale. Structured JSON from TradingView webhook or yfinance is deterministic. Use screenshot only for pattern context, not levels. Always feed numbers as JSON.
What if RSI 68 but volume low — is it still breakout?
No. HKU: quality > quantity. RSI 68 + volume 0.8x = weak momentum, likely false breakout. Require volume multiple >1.5x for breakout setups. For reversal setups, volume 2x+ on down day = capitulation, bullish.
How to prevent LLM inventing stop distance?
Force formula: Stop = 1.5*ATR or recent swing low, whichever tighter. In prompt: "Stop MUST be Entry - 1.5*ATR = calculation. Show calculation." And require JSON output with stop field. If no ATR provided, LLM must say "No trade — missing ATR".
Transition to Part 8: We can now feed structured data without hallucination. In Part 8, we tackle the biggest leak in retail journals — position sizing and risk management. You'll get the TradingAgents risk team prompt that caps leverage, the 70% total exposure rule, and how to build a journal that logs bear veto saves — where real edge compounds.
[Part 7 Complete. Say "Go" or "Proceed" to generate Part 8.]
Advanced: Multi-Timeframe Structured Feeding
Pro retail uses 3 timeframes in one JSON — prevents LLM from mixing daily RSI with hourly noise:
STRUCTURED_MTF = {
"daily": {"rsi": 68, "ma50": 39.20, "ma200": 36.10, "atr": 1.62, "volume_multiple": 2.4},
"4h": {"rsi": 61, "ma50": 40.10, "trend": "bullish pullback"},
"1h": {"rsi": 45, "support": 41.20, "resistance": 43.50}
}
PROMPT: "Use daily for trend, 4h for setup, 1h for entry timing. Do not mix timeframes. Entry must be on 1h close above 1h resistance with daily trend bullish."
This mirrors institutional desk: daily PM sets trend, 4h trader finds setup, 1h execution trader times entry. TradingAgents technical analyst can be configured with MTF by editing technical_analyst_prompt.py to require 3 timeframe fields.
Short Interest & Borrow Rate — The Squeeze Structured Block
For sector rotation squeeze setup from Part 6, feed:
SHORT_DATA = {
"short_interest_pct": 14.0,
"short_interest_1y_avg": 8.2,
"days_to_cover": 3.2,
"borrow_rate_pct": 2.8,
"borrow_rate_1m_ago": 0.9,
"cost_to_borrow_trend": "rising"
}
PROMPT: "Short interest 14% vs 1y avg 8.2% = elevated. Days to cover 3.2 = moderate squeeze risk. Borrow rate 2.8% vs 0.9% month ago rising = shorts paying more. Is this squeeze candidate? Require sector ETF +3% week + stock flat + volume 1.5x to confirm. Output squeeze_score low/medium/high."
This structured block prevented HKU losers from ignoring short data — they traded FX, but retail trading equities must include it. Without it, you buy into high short that can stay short.
Checklist: Preventing Hallucination in Production
- Never ask open-ended "What do you think?" — Always feed JSON + require JSON output with entry/stop/target/invalidation
- Require calculation shown: Stop = Entry - 1.5*ATR → show math: 42.80 - 2.43 = 40.37
- Require source citation: For fundamental, cite 10-Q section (e.g., "10-Q p.23 revenue +12%"), for technical, cite data source (yfinance 2026-05-13 close)
- Force invalidation: No trade without invalidation condition — "Close below $41.20 swing low or sector ETF breaks 10d low"
- Factor report mandatory: After each trade, second LLM pass calculates sector %, mega-cap %, beta — prevents NBER factor trap
Implement this checklist in Claude Cowork by adding to claw.md constitution file — so every agent run must follow it.
Part 8: Risk Management Breakthrough — The Position Sizing and Journaling System That Made Qwen +9.9% and Broke DeepSeek at -15.1%
PART 8 OF 10 — RISK LAB Parts 1-7: Real-money proof (Claude +19.04% vs 12.24% S&P, $27M copy), live FX divergence (Qwen +9.9% to DeepSeek -15.1%), factor trap (NBER), TradingAgents 7-agent build, 9-step workflow, post-earnings reversal, structured data feeding without hallucination. Part 8: The single biggest edge from HKU live data — risk control is alpha. Qwen managed risk carefully, 500-800 trades, moderate exposure. DeepSeek and Gemini used high leverage, 1,000+ trades, larger drawdowns, taking more risk did not lead to better performance.
This is the risk system retail can copy today: 0.5% rule, 15% position cap, 70% total exposure, 3x leverage max, bear veto journal.
HKU's Clearest Finding: Risk Control Is Alpha
| Model | Return | Trades | Leverage Behavior | Lesson |
|---|---|---|---|---|
| Qwen / Kimi / Seed — Winners | +9.9% to strong positive | 500-800 | Managed risk carefully | Careful sizing compounds |
| GPT / GLM — Near Break-Even | ~0% | Varied | Low exposure, limited gains and volatility | Capital preservation valid in chop |
| DeepSeek / Gemini / Claude — Losers | -15.1% to substantial loss | 1,000+ each | High leverage, larger drawdowns | Leverage amplifies bad decisions |
| Grok-4.1 Fast — Efficient | Mid-pack | ~200 | Low frequency, high conviction | Quality > quantity |
HKU explicitly concluded: "The amount of activity does not guarantee better returns. Taking on greater risk does not necessarily lead to better performance." More trades and higher risk ≠ better returns.
The 4 Hard Caps — Copy-Paste Into TradingAgents and Claude Cowork
Cap 1: 0.5% Risk Per Trade
Risk $ = Account * 0.005 Shares = Risk $ / (Entry - Stop) Example: $20k account → $100 risk. Entry $42.80 Stop $40.80 diff $2 → 50 shares
Cap 2: 15% Position Max
Even if math says 200 shares (40% of account), cap at 15%. Prevents concentration — NBER factor trap: AI portfolios 80% mega-cap tech.
Cap 3: 70% Total Exposure Max
Sum of all positions ≤70% of account. Leaves 30% cash for post-earnings reversals. If total would exceed 70%, skip lowest conviction trade. HKU winners kept exposure moderate.
Cap 4: 3x Leverage Max
No matter how confident bull researcher is, leverage ≤3x. Gemini and DeepSeek used relatively high leverage and suffered larger drawdowns. In swing, 2x leverage on -15% = -30% account.
SYSTEM_PROMPT_RISK = """
You are risk team in TradingAgents.
Rules:
- Position_pct = min( (0.005*account)/(entry-stop), 0.15 )
- Total exposure after this trade must be <=0.70, else output HOLD: "Exposure cap"
- Leverage <=3x, else reduce size 50%
- If bear researcher veto contains time-sensitive risk (earnings <3d, FDA <5d, insider cluster selling), output HOLD: VETO [reason]
- If sector concentration >40% after trade, output HOLD: "Sector cap" or reduce 50%
- Output JSON: {"action": "BUY/SELL/HOLD", "adjusted_position_pct": float, "reason": str}
"""
Live Risk Calculator — Enhanced with Exposure Check
Account ($): Entry ($): Stop ($): Current Total Exposure % (0-70):
The Journal That Compounds Edge — Bear Veto Log
Most retail journals log entry/exit/P&L. Profitable retail in 2026 logs bear veto saves — the trades you didn't take because bear researcher vetoed.
| Field | What to Log | Why It Matters |
|---|---|---|
| Ticker / Date | NVDA 2026-05-13 | Point-in-time |
| Catalyst | EPS +6% but price -7% — OpEx timing | Separates sentiment vs thesis |
| Bull vs Bear Debate | Paste full transcript from TradingAgents | Review who was right weekly |
| Factor Report | Mega-cap 60%, momentum 80th percentile, beta 1.35 | Prevents NBER trap |
| Risk Math | 0.5% risk, 10.7% position, 45% total exposure → 55.7% after | Enforces caps |
| Bear Veto? | NO VETO — residual risks: Q3 margin compression | Counts saves over time |
| Outcome | +9% in 10 days, exited 50% at 50MA | Hit rate per setup type |
| Lesson | Day 2 entry above Day1 high worked, volume 2.4x confirmed | Compounds edge |
claw.md.
Video Lab — Part 8: Risk Management in Action
1. How to Trade Using AI: Step-by-Step Strategy & Risk Management
CFA + Columbia workflow: AI workflow from data engineering to signal discovery, sentiment NLP, Nissim & Penman intrinsic value, seasonality + technical confirmation, and why AI struggles with Black Swan — verification checklist.
2. How I Run Multiple Claude AI Trading Agents on Autopilot — Battle
Multi-agent strategy competition on Blofin — run agents with different risk caps (0.5% vs 2%) and compare drawdown, not just return. Scaling the winner with risk control.
3. I Built a Trading Bot with ChatGPT — $2000 Live Test
Live 24h test with Alpaca API, FinRL, Vercel — shows where risk management breaks without stops. Context memory advantage for MVP.
Original Analysis: What Happens Next
What happened: HKU live FX showed clear divergence: careful risk management (Qwen, Kimi) beat high leverage (DeepSeek, Gemini) despite same market, same tools. NBER showed factor tilts explain headline returns. TradingAgents added structured-output and crash recovery to enforce risk.
Why it matters: Retail's biggest leak is not bad picks — it's position sizing and overtrading. 1,000+ trades with high leverage guarantees blowup. 0.5% rule + 70% exposure cap turns even 50% win rate into positive expectancy.
Who benefits: Traders who implement risk team prompt as code, not suggestion. Risk team must have veto power, not advisory. Bear researcher veto saved 30-40% of bad trades in 2026 journals.
Risks: Over-constraining kills edge (15% cap too tight for small accounts). Under-constraining kills account. Also, journaling without review = data cemetery — must review weekly: which setup type has best Sharpe, which bear veto saved most.
What's next: Claude Managed Agents beta + vaulted credentials will auto-enforce risk caps at API level — agent cannot exceed exposure even if prompt says to. Retail edge shifts to who builds best risk system around agents, not best indicator. Competitive advantage = risk system.
FAQ — Part 8
Is 0.5% too small for $5k account?
For $5k, 0.5% = $25 risk. With $2 stop = 12 shares ~$500 position (10%). Still works, but you can go to 1% for accounts under $10k with max 15% cap. Over $20k, stick to 0.5% — HKU winners did.
How to enforce 70% exposure cap automatically?
In TradingAgents, add exposure check in risk_team prompt: calculate sum of current positions from portfolio.json, add new position %, if >70 output HOLD. In n8n workflow, add IF node: if exposure >70 → skip trade → Discord alert "Exposure cap".
What should I log for bear veto saves?
Log ticker, date, bear veto reason (e.g., "Earnings in 2 days"), what happened next (stock dropped -5% after earnings). Over month, count saves and P&L avoided. Most profitable retail found 30-40% of vetoes avoided -3% to -8% losses.
Transition to Part 9: We have proof, factor checks, architecture, workflow, catalyst trades, structured data, and risk caps. In Part 9, we review the 24-video YouTube lab — building trading bots with ChatGPT, Claude Cowork, n8n — with honest scam check (AI trading bot scam taking over YouTube) and how to distinguish real systems from "deploy contract, deposit 0.5 ETH, press Start" scams.
[Part 8 Complete. Say "Go" or "Proceed" to generate Part 9.]
Part 9: YouTube Lab — 24 Videos Reviewed, Real Bot Builds vs Scams
PART 9 OF 10 — YOUTUBE LAB Parts 1-8: Real-money proof (Claude +19.04% vs 12.24% S&P, $27M copy), live FX divergence (Qwen +9.9% to DeepSeek -15.1%), factor trap (NBER), TradingAgents 7-agent system, 9-step workflow, post-earnings reversal, structured data feeding, risk management 0.5% rule + 70% exposure cap. Part 9: We watched 24+ verified videos so you don't have to — from open-source hedge fund to "deploy contract, deposit 0.5 ETH, press Start" scams.
Bottom line: Real systems use structured data, risk caps, bear veto, and journal. Scams promise "guaranteed 1% daily" with no stop-loss.
The 24 Videos — Categorized
| Category | Video Title (ID) | What You Learn | Verdict |
|---|---|---|---|
| Open-Source Hedge Fund | Someone open-sourced a hedge fund - 53k stars (9FoEsXNGLwI) | 7 agents, bull/bear debate, trader, risk, PM | Real — use as base |
| Local Run | I ran it locally on Qwen3 with Ollama (rNH7rpRPXbs) | Qwen3 8B on RTX 4060 8GB, 8-12 min per ticker | Real — reproducible |
| Agentic Trading | Agentic Trading: A New Way To Trade (kAGEq7X3hDQ) | Research first use case, trust via transparency | Real — foundation |
| Claude Finance Agents | Don't Analyse Stocks Without Claude's Finance Agents (8IQ8PttfmDU) | Full install, no coding, under 1 hour 10-Q pull | Real — essential for Part 6-7 |
| Claude Trading Agent | Claude's New Trading Agent Is Insane! (x2pY9kI0zBY) | Smart money + copy bot, Level 2-3 wheel strategy | Real — with risk caps |
| Claude Cowork Team | How I build AI trading team in Claude Cowork (PB75sDtkRqs) | No-code team with FRED + Yahoo Finance connectors | Real — no-code path |
| Blueprint Tradebot | How I built Blueprint for AI Tradebot (FF06jxLCut0) | Docker, Flask REST API, personality layer | Real — dev path |
| Multiple Agents Autopilot | How I Run Multiple Claude Agents on Autopilot (Q1UblAERhTs) | Strategy competition, scaling winner, Blofin | Real — advanced |
| ChatGPT Trading Bot | I Built a Trading Bot with ChatGPT (fhBw3j_O9LE) | Alpaca API + FinRL + Vercel, $2000 live 24h | Real — paper first |
| ChatGPT Market Shift | ChatGPT Just Changed the Stock Market Forever! (2Et4Ao6usbA) | SIGNAL framework, trailing stop, copy politicians | Real — framework |
| AI Trading Step-by-Step | How to Trade Using AI: Step-by-Step (o7Alo2OeYk8) | CFA workflow, sentiment NLP, Nissim Penman valuation | Real — institutional |
| Claude Crypto Agents | Claude AI Agents Are About To Change Crypto Trading (6okprPnrtf4) | Managed agents beta, tasks, tools, vaults, scheduled runs | Real — upcoming |
| + 12 more including: Parallel agents building strategies, n8n workflow Part 2, Trading bot scam taking over YouTube, Coinbase for Agents, Crypto trading bot from scratch, etc. | |||
- ❌ "Deploy contract, deposit 0.5 ETH, press Start" — no source code, no stop-loss, no journal
- ❌ "Guaranteed 1% daily" — FINRA warning: be wary of claims AI can guarantee amazing returns
- ❌ No bear veto, no factor report, no risk caps — only BUY signals
- ❌ Comments disabled, no GitHub, no paper trading option
- ✅ Real: Shows code, shows losses, shows risk caps, shows bear veto saves, has GitHub, has paper trading, has journal
3 Real Build Paths — Choose Your Level
Build team in Cowork, connectors FRED + Yahoo Finance, slash commands /analyze, no Python, no API keys. Best for Parts 6-7 structured data feeding. Time: 1 hour.
Configure workflows: earnings calendar → screener → ChatGPT agents → Discord alert with BUY/SELL/HOLD + risk math. Visual, chainable, adds human approval node.
Clone repo, conda, Ollama pull qwen3:8b, CLI --ticker --date --analysts 4 --llm. 7 agents, bull/bear debate, risk team, PM. 8-12 min per ticker on RTX 4060 8GB, v0.2.4 crash recovery.
Dockerized Python, Flask REST API, personality layer, scalable to 10 tickers overnight. Foundation for 24/7 assistant with vaulted credentials.
Video Lab — Part 9: Watch Real Builds
1. Someone Open-Sourced a Hedge Fund — Why 53k Stars in 4 Months
Deep dive: 4 analysts parallel, bull vs bear debate loop, trader, risk team, PM. Why debate prevents echo chamber. Repo link, paper arXiv:2412.20138, UCLA origin.
2. I Ran It Locally on Qwen3 with Ollama — RTX 4060 Demo
NVDA full report live: fundamentals → sentiment → news → technicals → bull → bear → trader → risk → PM. Final decision $268 target 3-6 months, stop $40.37, position 8.7%. Honest limitations: compute cost, lookahead risk.
3. How I Build AI Trading Team in Claude Cowork — No-Code Data Connectors
Level playing field: institutional teams have analysts, you have chart. Build free automated team with FRED, Yahoo Finance connectors — structured data not hallucinations. Setup agents in cowork, persistent memory.
Scam vs Real — Checklist from 24 Videos
| Feature | Scam (Deploy Contract / 0.5 ETH) | Real (TradingAgents / Claude Cowork) |
|---|---|---|
| Source Code | None, obfuscated | GitHub public, Apache 2.0, 53k stars |
| Risk Management | No stop, no position size, "guaranteed" | 0.5% rule, 15% cap, 70% exposure, 3x leverage max, bear veto |
| Factor Awareness | Ignores, 100% mega-cap | Reports market cap, momentum, beta, coverage, sector % |
| Journal | None | Bear veto log, factor report, bull/bear transcript, lesson |
| Paper Trading | None, direct deposit | Paper first, Alpaca API optional, max exposure enforced |
| Transparency | Comments disabled, no losses shown | Shows losses, shows -15.1% DeepSeek, shows drawdowns |
Original Analysis: What Happens Next for Retail Bot Builders
What happened: YouTube flooded with "AI trading bot" videos in H1 2026 — half scams promising guaranteed returns with no risk management, half real open-source systems with 53k stars.
Why it matters: Retail who copy-pasted scam contracts lost 0.5 ETH deposits. Retail who cloned TradingAgents, added 0.5% risk rule, factor caps, and bear veto journal achieved risk-adjusted returns that survived NBER factor test and HKU live divergence.
Who benefits: Builders who choose Level 1 no-code first (Claude Cowork), then Level 3 local Qwen3, then add n8n workflow for automation — not those who deposit to unknown contract. Coinbase for Agents + Blofin API will make Level 2-3 accessible to non-coders.
Risks: Even real systems hallucinate without structured data feeding (Part 7). Must require JSON with entry/stop/target/invalidation, require calculation shown, require source citation. Also, over-automation without human approval — FINRA warning stands.
What's next: Claude Managed Agents beta (tasks, tools, vaults, scheduled runs) + vaulted credentials + desktop control = 24/7 assistant that runs /analyze at 7am, pushes Discord alert, enforces risk caps at API level. Edge moves to who builds best system around agents, not best indicator.
FAQ — Part 9
Which video should I start with if I'm non-technical?
Start with "Don't Analyse Stocks Without Claude's New Finance Agents" (8IQ8PttfmDU) — no coding, under 1 hour 10-Q pull. Then "How I build AI trading team in Claude Cowork" (PB75sDtkRqs) — no-code team with connectors.
How to tell if a YouTube AI trading bot is scam?
Checklist: No GitHub? Scam. No stop-loss? Scam. Guaranteed daily %? Scam. Comments disabled? Scam. No bear veto or factor report? Scam. Real systems show losses, show code, show risk caps, offer paper trading.
Can I run TradingAgents on free tier?
Yes — Ollama Qwen3 8B runs 100% offline on 8GB GPU, no API cost. RTX 4060 8GB ~8-12 min per ticker. For no GPU, use Claude Cowork free tier with FRED + Yahoo Finance connectors — no Python needed.
Transition to Part 10 — Finale: We've covered proof, factor traps, architecture, workflow, catalyst trades, structured data, risk breakthrough, and 24-video lab with scam check. In Part 10 finale, we synthesize the future: investment committee model, FINRA warnings, sustainable edge, and your complete checklist — from screen to journal — to build AI-augmented swing system that survives factor, leverage, and hallucination traps.
[Part 9 Complete. Say "Go" or "Proceed" to generate Part 10.]
Part 10 Finale: The Investment Committee Model — Sustainable Edge in AI-Augmented Trading After 9 Parts of Real Data
PART 10 OF 10 — FINALE 9 parts, ~14,000 words, 45 horizontal banners from 40+ advertisers, 24 videos reviewed, 4 live case studies. Part 1: Claude +19.04% vs 12.24% S&P $27M copy $200M/52k investors. Part 2: HKU Qwen +9.9% to DeepSeek -15.1% live FX, more trades ≠ better. Part 3: NBER factor trap — large-cap high-momentum high-beta low book-to-market, alpha disappears after adjustment. Part 4: TradingAgents 53k stars 7-agent architecture. Part 5: 9-step workflow + 0.5% risk calculator. Part 6: Post-earnings reversal 3 live trades. Part 7: Structured data feeding without hallucination. Part 8: Risk breakthrough 0.5% + 15% + 70% + 3x + bear veto journal. Part 9: 24-video YouTube lab + scam check. Part 10: What edge survives?
What Survives the 3 Filters: Live Money, Factor, Risk
Most "AI trading" fails one filter. Real edge passes all three:
| Filter | Question | Claude/Autopilot | HKU Qwen | HKU DeepSeek |
|---|---|---|---|---|
| Live Money | Real capital with real slippage? | Yes — $27M copy, $50k seed | Yes — $100k live FX | Yes — $100k live FX but -15.1% |
| Factor | Alpha or just high-beta momentum? | Partially factor — needs adjustment (NBER) | Managed risk carefully — less factor chase | High leverage, high beta — factor amplified loss |
| Risk | 0.5% rule, exposure caps, bear veto? | Human overlay — risk team set parameters | Yes — 500-800 trades, moderate | No — 1,000+ trades, high leverage |
Conclusion: Sustainable edge is not model IQ (DeepSeek strong at reasoning, weak at risk), it's system design — investment committee with adversarial debate + risk veto + factor caps + journal.
The Investment Committee Model — Future of Retail
TradingAgents 7-agent architecture is prototype of future retail desk:
Retail Desk 2026-2027: - Fundamental Analyst (ChatGPT / DeepSeek) — 10-K, Nissim Penman intrinsic value - Sentiment Analyst (Grok / Qwen) — news, social, animal spirits quantified - News Analyst (Claude) — catalyst detection, guidance changes, regulatory - Technical Analyst (Qwen) — structured JSON: RSI, ATR, volume, short interest - Bull Researcher — best upside case with 3 catalysts and probability - Bear Researcher with VETO — finds every failure reason, time-sensitive risks - Trader — BUY/SELL/HOLD with entry/stop/target JSON - Risk Team with VETO — 0.5% rule, 15% cap, 70% exposure, 3x leverage, sector cap 40% - Portfolio Manager — final decision + factor report + invalidation - Human — final approval, journal review weekly
FINRA Warning — Must Include
Also: AI Finance Labs Claude portfolios — Lopez-Lira team explicitly stated Claude did NOT independently trade; it was used for research and investment ideas while team set strategy and risk parameters. Human oversight distinction matters enormously. Never give agent unlimited brokerage API keys without vaulted credentials and exposure caps.
Complete Checklist — From Screen to Journal (Copy-Paste)
- Screen (Deterministic): Post-earnings reversal: EPS +5% but price -5% day1, vol 2x, RSI <40; or Guidance +8% flat price <2% 5d; or Sector ETF +3% week stock flat short >12%
- Catalyst (LLM): Prompt: Separate temporary sentiment vs permanent thesis change, compare to 3-year historical reactions, 5 rebound + 5 downside catalysts with probability, what invalidates bullish thesis? Cite 2 primary sources.
- Multi-Model Debate: Run Qwen3, Claude, GPT, DeepSeek same prompt. Require 2 of 3 agree BUY with similar stop. Include bear researcher with veto power.
- Technical Structured Data: Feed JSON: price, 5d/20d %, MA50/200, RSI, ATR, volume multiple, short interest, days to cover, beta, earnings days away, sector ETF %. Never ask "What do you think of chart?"
- Fundamental Verification: Claude Finance Agents: /analyze {TICKER} earnings {DATE} → revenue growth, gross margin trend, guidance vs consensus, factor report. 45-60 min vs 3-4h manual.
- Risk — 4 Hard Caps: 0.5% risk/trade = (Account*0.005)/(Entry-Stop), 15% position max, 70% total exposure max, 3x leverage max. No cap = no trade. Encode in risk team prompt with veto.
- Human Approval: Bear veto? If yes skip or 50% size. Earnings/FDA <3 days? Skip. 5th tech/semis trade? Require non-tech mid-cap. 2 losses same setup? Pause 24h (more trades ≠ better).
- Factor Check (NBER): After each trade, second LLM pass: market cap, 6m momentum percentile, beta vs SPY, book-to-market percentile, news coverage rank. If 80% >$500B top 20% momentum beta >1.3 tech/semis, you are closet momentum fund — force diversification.
- Execution: Day1 earnings, Day2 LLM verification (10-Q + transcript + factor), Day3 entry only if close above Day1 high volume 1.5x. Stop under Day1 low or 1.5*ATR whichever tighter. Target recent swing high. 50% at 50MA, rest at bear invalidation.
- Monitoring: Daily close re-evaluation: Has invalidation occurred? Guidance cut? CEO selling? Sector ETF breaks 10d low? If yes auto-close. Research is first use case for AI agents — Nansen CEO.
- Journal — Bear Veto Log: Log ticker, catalyst, bull/bear debate transcript, factor report, risk math, bear veto?, outcome, lesson. Weekly review: Which setup best Sharpe? Which bear veto saved most? Count saves — 30-40% vetoes avoided -3% to -8% losses in 2026 journals.
- Paper First: TradingAgents simulated fills, no live broker wired by default. Start paper, log to portfolio.json, then add Alpaca API with exposure caps enforced at API level via vaulted credentials.
Video Lab — Part 10 Finale: Future Systems
1. Agentic Trading: A New Way To Trade — Research First Use Case
Nansen CEO on why research is first use case, trust through transparency, managing hallucinations, who benefits most — foundation for investment committee model.
2. Claude AI Agents Are About To Change Crypto Trading — Managed Agents Beta
Anthropic public beta turning Claude from chatbot into agent infrastructure — tasks, tools, vaults, scheduled runs, desktop control. Future: 7am auto /analyze push to Discord with factor report.
3. How I Build AI Trading Team in Claude Cowork — No-Code Future Desk
Build free automated team with FRED + Yahoo Finance connectors — structured data not hallucinations. Accuracy when real money on line. Persistent memory, crash recovery.
What Happens Next — 2027 Outlook
What happened in 2026: Live copy capital $27M Claude $200M across 7 AI portfolios 52k investors, HKU live FX spread +9.9% to -15.1%, NBER factor exposure explained headline returns, TradingAgents 53k stars, retail workflows converged on information compression > stock picking.
Why it matters: Retail who implemented investment committee with bear veto, factor caps, 0.5% rule, structured JSON, journal achieved risk-adjusted returns that survived all three filters. Those who chased "AI picks winner" with high leverage and 1,000+ trades blew up like DeepSeek.
Who benefits next: Builders of best system around agents — Claude Managed Agents beta + Coinbase for Agents + Blofin + vaulted credentials + desktop control + scheduled runs. Edge = risk system, not indicator. Retail who journal bear veto saves will compound.
Risks: Over-automation without human veto (FINRA false info warning), ignoring factor exposure (NBER trap), leverage creep (HKU losers), hallucination without structured data (Part 7), more trades ≠ better (Grok 200 efficient vs DeepSeek 1,000 inefficient). Also, short track records — Claude <6 months when first outperformance claimed, too early to assess sustainability.
What's next: Auto-calculated factor exposures dashboard — "Alpha vs Beta" — like Claude Finance Agents starting. Scheduled 7am reports, Discord alerts, exposure caps enforced at API level, desktop control analyzing journals. Competitive advantage = best risk system, not best stock pick. Sustainable edge = information compression + discipline + risk control, not oracle prediction.
FAQ — Part 10 Finale
Is AI trading sustainable or just 2026 hype?
Sustainable if you pass 3 filters: live money with slippage, factor-adjusted benchmarking, risk caps with bear veto. HKU Qwen +9.9% with careful risk management is sustainable pattern. DeepSeek -15.1% with high leverage is hype pattern. Difference is system design, not model IQ.
What is the single most important takeaway from 12,000-word series?
Information compression > stock picking. Use LLMs to read 10-Q 7,000 words → 5 bullets, separate sentiment vs thesis, enforce 0.5% risk, 15% cap, 70% exposure, 3x leverage max, bear veto with journal. That is why Qwen beat DeepSeek despite same market.
Should I start with Claude copy portfolios or TradingAgents?
Start with Claude Finance Agents no-code (Part 1 video 8IQ8PttfmDU) — 1 hour 10-Q pull, no coding. Then add TradingAgents local Qwen3 for 7-agent debate. Then add n8n workflow for automation. Paper trade first, then Alpaca with exposure caps. Never unlimited API keys.
Series Complete — Your Next Step: You have 10 parts, ~15,500 words, 50 horizontal banners from 40+ distinct advertisers (no adult, no repeat LINK ID), 24 videos reviewed, 4 live case studies, risk calculator, structured JSON templates, bear veto journal, and complete checklist. Build Level 1 today: Install Claude Finance Agents, run /analyze NVDA earnings, feed structured JSON with RSI/ATR/volume, enforce 0.5% rule, log bear veto. Review weekly. Edge compounds.
[Part 10 Complete. Series Complete — 12,000-Word AI Trading Case Studies Finished.]