Trading Strategies & Tools · 11 min read · May 23, 2026

Opus 4.7 vs Gemini 3.1: What Running Both on a Live Forex EA Actually Showed Me

Most "AI trading model comparisons" you'll find online were written by someone who tested each model for ten minutes on a backtest. I've been running Claude Opus 4.7 and Gemini 3.1 Pro on the same live XAUUSD account for the last couple of months. Same prompts, same risk, same broker.…

Most “AI trading model comparisons” you’ll find online were written by someone who tested each model for ten minutes on a backtest. I’ve been running Claude Opus 4.7 and Gemini 3.1 Pro on the same live XAUUSD account for the last couple of months. Same prompts, same risk, same broker. This is what actually showed up.

Why this comparison is different from the ones you’ve already read

Open YouTube right now and search “Opus 4.7 vs Gemini 3.1.” You’ll get reasoning benchmarks. Coding benchmarks. Math benchmarks. Tons of side-by-side prompts with cherry-picked answers.

None of that tells you what these models do when there is real money on the screen and the next decision could lose €200.

That is a completely different problem from “write me a poem about Tuesday.” It involves uncertainty, partial information, asymmetric risk, and a hard time limit. Two models that score identically on a public benchmark can behave very differently when you wire them into a forex EA and let them decide whether to take a XAUUSD setup at 14:32 on a Friday with US news in 18 minutes.

I run both of them, on the same live forward-test account, every week. The decisions are public on the Myfxbook. The video sessions are public on my YouTube channel. So instead of giving you another benchmark recap, I’m going to walk you through the patterns I keep seeing — the ones nobody talks about because nobody actually runs both with real fills.

If you’re considering an AI-driven EA and you want to know which model to use, this should save you a lot of API budget.

Setup: how I run Opus 4.7 and Gemini 3.1 Pro on the same EA

Before the comparison makes any sense, the methodology matters. Otherwise you get a “Claude is smarter” or “Gemini is faster” take that depends entirely on the prompt the writer happened to throw at it.

Here is how the test runs:

  • Same EA: DoIt Alpha Pulse AI — the LLM-decision-layer EA I built. It accepts any model with an API key. You swap providers in the panel; the rest of the system stays identical.
  • Same instrument: XAUUSD on M15.
  • Same broker, same account: live forward test on Axi Select, RoboForex setup variant for cross-validation.
  • Same risk per trade: ~0.5% on the live forward, lower than I’d run on a personal account.
  • Same prompt structure: the system prompt that ships with Alpha Pulse, with the same context window (price action levels, news pulse, recent volatility, account state).
  • What changes: only the model. Opus 4.7 one week, Gemini 3.1 Pro the next, occasionally side-by-side on alternating signals.

If you want to see the live decisions instead of reading my summary, the Myfxbook is open and the live XAUUSD sessions are on YouTube every week. The Alpha Pulse landing has both linked from the live results section.

Three lessons from this experiment, regardless of what you buy

  1. Public benchmarks don’t transfer to live trading. The model that wins on coding tasks isn’t necessarily the one that wins on price decisions. Test what you actually care about, with money on the screen, before committing.
  2. Cost-per-decision is the metric most people forget. A model that is 5% smarter but 3x more expensive is not a better trading model. Run the math on your decision frequency before you choose.
  3. Multi-model voting beats single-model optimization. If you have the budget, running two models and taking only the trades they agree on is the highest-conviction setup I’ve found. The disagreement itself is information.

If you read this far and you’re still on the fence about whether AI-driven EAs are worth the effort: the honest answer is that they are, but not for the reason most YouTubers tell you. They don’t replace risk management. They don’t predict the market. What they do is make decisions that are slightly more contextual than a static rule-based EA could ever make. Over a year, “slightly more contextual” compounds.

FAQ

Is Opus 4.7 the same model people call “Claude Opus 4.7”?

Yes. Anthropic’s flagship reasoning model, released in 2026. The full name is Claude Opus 4.7. I use the short form throughout this post for readability.

Is Gemini 3.1 Pro the same as Gemini 3 Pro?

No. Gemini 3.1 Pro is the updated version released in 2026, more capable on structured reasoning and instruction-following than the original 3.0 series. If you’re testing AI EAs and your provider only supports 3.0, the experience will be noticeably weaker.

Can I really run both models on the same EA?

Yes, that’s the design of DoIt Alpha Pulse AI. You bring API keys for Anthropic, Google, OpenAI, xAI, DeepSeek or Qwen and switch in the panel. There’s no model lock-in. More setup questions answered here.

What about GPT-5.5 and Grok 4.20?

Both supported in Alpha Pulse. I tested them less rigorously in this round because Opus and Gemini are the two models my live forward test runs primarily on. GPT-5.5 sits between Opus and Gemini in my experience — slightly better than Gemini on reasoning, slightly worse than Opus on confidence calibration. Grok is wildly inconsistent in trading contexts; useful as a third opinion in a voting setup, not as a primary.

How much does running this cost in API fees?

Depends on the model. Gemini 3.1 Pro on a typical week: $15-25. Opus 4.7 the same week: $40-70. Running both for a voting layer roughly doubles. None of these costs include trading capital, slippage, or commission — those are separate.

Where can I see the live results?

The DoIt Alpha Pulse AI live forward test account is on Myfxbook, public. The link is on the product page under live results. Every winner, every loser, every drawdown period — visible without having to take my word for any of it.


Trading AI EAs involves real risk to capital. The performance described is from a live forward-test account on Axi Select / RoboForex. Past results don’t guarantee future ones. Run on demo first; size at low risk; treat any AI as a decision-support layer, not a guarantee.

Want updates when the multi-model voting layer ships? The newsletter is where it gets announced first: subscribe here.

Diego Arribas
Diego Arribas
Founder · DoItTrading

Building MT4/MT5 expert advisors and writing about prop-firm scaling since 2021. Currently running Alpha Pulse AI live on XAUUSD and trading Axi Select in parallel. I write what I'd want to read before paying for any of this myself.

Scroll to Top