From hypothesis to evidence: autonomous strategy research with MCP
Use an AI assistant and Stratifyre MCP to preserve rules, backtest IDs, trade evidence, and rejected hypotheses in one bounded research workflow.
An assistant’s explanation of a trading idea is easy to produce. A saved strategy, completed backtest, and traceable trade require a different standard: evidence that the assistant actually used the account and tools you authorized.
This walkthrough records an assistant-directed Stratifyre MCP experiment: two candidate strategies, three completed historical backtests, a losing trade inspection, and one refinement tested on a reserved period. It also records a failed creation call. The outcome was useful research, including a rejected filter and a losing refinement; it did not establish a trading edge.
Give the assistant a bounded research task
MCP connects an AI application to a server’s tools and resources; the application directs the workflow. The protocol itself does not decide which hypotheses to test. Official MCP architecture
If your goal is to have Claude or another assistant backtest a trading strategy, start by specifying the experiment and permissions. “Find a profitable Bitcoin strategy” leaves the search budget, execution assumptions, and stopping rule open.
Here, the question was narrower: does adding an hourly SMA50 trend filter change a fixed SMA10/20 crossover’s trades and drawdown? Two development runs answered that comparison. A third evaluated one subsequent change on later dates.
- Question and permissions → resolved instrument and fixed assumptions
- Live MCP discovery → validated payload → saved strategy readback
- Backtest ID → terminal status → metrics and every returned trade
- Observed weakness → one frozen refinement → reserved-period result
Author-created process map; each arrow requires its own evidence.
Verify the actual connection first
The recorded client was an isolated instance of the official TypeScript MCP SDK, version 1.28.0, driven by this assistant. It used Streamable HTTP and a normal browser OAuth authorization-code flow with PKCE. This was an actual MCP interaction, not a Claude or ChatGPT interface demonstration.
Before writing, the MCP user-profile identity was compared privately with the authorized demo browser account. The live connection exposed 26 tools under six granted scopes: strategy and backtest read/write, plus user and billing read. That count describes this connection’s permissions, not a permanent server-wide catalogue.
Discover the live schemas and begin with an account read. Treat tool availability, account ownership, and permission to submit a billable job as separate checks. Stratifyre’s documented assistant workflow calls for reviewing mutation payloads, preserving returned IDs, and checking saved state after interruptions. Conversation approval is workflow guidance, not proof of a server-enforced approval screen.
Fix the comparison before viewing results
The instrument was resolved through the product’s instrument index before MCP research: Binance BTC/USDT spot, canonical ID CRYPTO:BINANCE:BTCUSDT. Instrument search was not performed through MCP.
| Assumption | Recorded specification |
|---|---|
| Development dates | January 13–19, 2025, inclusive date request |
| Interval and market | Hourly close-source indicators; long-only spot; UTC; no leverage |
| Candidate A entry | SMA10 crosses above SMA20, position quantity zero, pending buy quantity zero |
| Candidate B entry | A’s entry plus close strictly above SMA50 |
| Exit | Flatten when SMA10 crosses below SMA20 while long; also flatten at the run’s end |
| Size and cash | Fixed 0.01 BTC; $10,000 initial cash; one position |
| Execution inputs | on-close, pessimistic fills; no volume cap or market impact |
| Cost inputs | Zero commission; assumed 0.05% adverse slippage, or five basis points |
Both saved backtest configurations matched these inputs. Zero commission deliberately excludes exchange fees; these figures are not verified after-venue-fee returns. Spot results also say nothing about funding, borrowing, or liquidation.
The protocol requested prewindow indicator history, but the live MCP schema offered no warmup field and the early snapshot query returned no records. Exact automatic warmup and complete source-bar coverage remain unverified. The saved on-close setting also does not independently establish causal signal-to-fill timing. These runs demonstrate an inspectable workflow under recorded modeling assumptions, rather than validate those assumptions for trading.
Audit the saved rules, including failures
The first actual strategy_create call failed with column StrategyCreationOperation.recoveryOwnerToken does not exist. No successful natural-language conversion is claimed.
The supported alternative exposed by the same live manifest was to clone an owned demo strategy, replace all copied rules through strategy_update, and read the result with strategy_get. The original stayed unchanged. The assistant authored A and B explicitly; strategy_validate_rules returned two valid rules and no errors for each before the updates.
Cloning produced an ID, not evidence that the intended rules ran. The research record kept each candidate’s validated payload, persisted rule readback, strategy ID, backtest ID, and terminal status separately. Copied baseline statistics were excluded from the comparison.
Compare completed jobs and inspect the loss
Both development jobs reached complete. Their full returned trade lists reconciled to reported P&L within floating-point precision.
| Development candidate | Trades | Reported P&L | Reported maximum drawdown |
|---|---|---|---|
| A: SMA10/20 crossover | 3 | $7.04 | 0.8567% |
| B: A plus SMA50 filter | 2 | −$30.20 | 0.8599% |
The trade list explained the difference. B excluded A’s $37.23 winner while keeping its $20.19 winner and $50.39 loss. The filter did not reduce reported drawdown here; three trades cannot establish how it behaves generally.
The selected loss entered at $104,539.173465 and exited at $99,500.225, with quantity 0.01 BTC and reported commission zero:
(99,500.225 − 104,539.173465) × 0.01 = −50.38948465
The matching native trade record spans January 18 at 19:00 UTC to January 20 at 00:00 UTC, the run boundary. End-of-run flattening was enabled, so this record does not prove the loss closed through the crossover exit. It also prompted a one-day buffer before the reserved evaluation, rather than assuming adjacent date labels could not overlap.
Freeze one refinement, then keep its losing result
A’s largest loss was about 4.82% of the position’s entry value. The assistant retained A as the base and proposed one change: an entry-owned initial stop at order.entry_price * 0.97, anchored to the actual first fill. The original A and B stayed unchanged.
The updated payload was validated and re-read before the reserved January 21–27, 2025 run. That later period had not been inspected during selection. The third job completed with seven losing trades, −$95.70 P&L, and 0.9947% reported maximum drawdown.
One returned trade lost about 3.0485% of its entry value. A 3% stop trigger therefore did not imply a 3% realized-loss cap in this run. More generally, a triggered stop becomes a market order rather than guaranteeing its trigger price. Investor.gov stop-order explanation
The research stopped after the three prespecified jobs. There was no additional parameter search to replace the losing result. A larger validation program would need untouched data, verified timing and costs, and explicit selection rules; the separate walk-forward example explains that next research stage.
Reuse the evidence standard
Start with a request like this, supplying your own resolved instrument and windows:
Research one hypothesis using my authorized Stratifyre account.First discover live MCP tools, schemas, resources, and scopes;verify the intended account and read available limits.
Draft two labeled candidates for the supplied canonical instrument.Fix dates, warmup, signal/fill timing, capital, sizing, and costs.Reserve a later period before inspecting results.Show each complete write/job payload and its consequences for approval.
Use at most one active backtest and three jobs total.Read saved rules before submitting each job; poll terminal states.Keep every failure, empty result, loss, strategy ID, and job ID.Inspect at least one trade and reconcile quantity, prices, fees, and P&L.
Propose one evidence-motivated refinement; freeze it before the reserve.Separate observed results from unresolved assumptions.Do not submit live orders, start scanners, or contact recipients.The reusable output is the candidate-to-result notebook, including rejected ideas and unresolved checks. After this experiment, the owned OAuth tokens were revoked and a request using the revoked access token returned HTTP 401; completed strategies and reports remain available for inspection.
Use the AI assistant connection guides to establish your own supported client, then begin with one bounded question and the same evidence trail.
Put your strategy rules to the test
Build your strategy, inspect historical trades, and review the assumptions behind your results.
Build and test your strategy

