← All articles

From hypothesis to evidence: autonomous strategy research with MCP

Use an AI assistant and Stratifyre MCP to preserve rules, backtest IDs, trade evidence, and rejected hypotheses in one bounded research workflow.

AutomationPublished By Stratifyre

Topics

MCPAI AssistantsStrategy ResearchBacktesting

An assistant’s explanation of a trading idea is easy to produce. A saved strategy, completed backtest, and traceable trade require a different standard: evidence that the assistant actually used the account and tools you authorized.

This walkthrough records an assistant-directed Stratifyre MCP experiment: two candidate strategies, three completed historical backtests, a losing trade inspection, and one refinement tested on a reserved period. It also records a failed creation call. The outcome was useful research, including a rejected filter and a losing refinement; it did not establish a trading edge.

Give the assistant a bounded research task

MCP connects an AI application to a server’s tools and resources; the application directs the workflow. The protocol itself does not decide which hypotheses to test. Official MCP architecture

If your goal is to have Claude or another assistant backtest a trading strategy, start by specifying the experiment and permissions. “Find a profitable Bitcoin strategy” leaves the search budget, execution assumptions, and stopping rule open.

Here, the question was narrower: does adding an hourly SMA50 trend filter change a fixed SMA10/20 crossover’s trades and drawdown? Two development runs answered that comparison. A third evaluated one subsequent change on later dates.

Research record
  1. Question and permissions → resolved instrument and fixed assumptions
  2. Live MCP discovery → validated payload → saved strategy readback
  3. Backtest ID → terminal status → metrics and every returned trade
  4. Observed weakness → one frozen refinement → reserved-period result

Author-created process map; each arrow requires its own evidence.

Verify the actual connection first

The recorded client was an isolated instance of the official TypeScript MCP SDK, version 1.28.0, driven by this assistant. It used Streamable HTTP and a normal browser OAuth authorization-code flow with PKCE. This was an actual MCP interaction, not a Claude or ChatGPT interface demonstration.

Before writing, the MCP user-profile identity was compared privately with the authorized demo browser account. The live connection exposed 26 tools under six granted scopes: strategy and backtest read/write, plus user and billing read. That count describes this connection’s permissions, not a permanent server-wide catalogue.

Actual Stratifyre OAuth consent permissions for strategy and backtest read/write, user read, and billing read
The real consent screen's requested-permissions panel. The signed-in account panel is excluded. No scanner or live-trading scope was requested; granting a write scope remains broader than this experiment's instructions.

Discover the live schemas and begin with an account read. Treat tool availability, account ownership, and permission to submit a billable job as separate checks. Stratifyre’s documented assistant workflow calls for reviewing mutation payloads, preserving returned IDs, and checking saved state after interruptions. Conversation approval is workflow guidance, not proof of a server-enforced approval screen.

Fix the comparison before viewing results

The instrument was resolved through the product’s instrument index before MCP research: Binance BTC/USDT spot, canonical ID CRYPTO:BINANCE:BTCUSDT. Instrument search was not performed through MCP.

Assumption Recorded specification
Development dates January 13–19, 2025, inclusive date request
Interval and market Hourly close-source indicators; long-only spot; UTC; no leverage
Candidate A entry SMA10 crosses above SMA20, position quantity zero, pending buy quantity zero
Candidate B entry A’s entry plus close strictly above SMA50
Exit Flatten when SMA10 crosses below SMA20 while long; also flatten at the run’s end
Size and cash Fixed 0.01 BTC; $10,000 initial cash; one position
Execution inputs on-close, pessimistic fills; no volume cap or market impact
Cost inputs Zero commission; assumed 0.05% adverse slippage, or five basis points

Both saved backtest configurations matched these inputs. Zero commission deliberately excludes exchange fees; these figures are not verified after-venue-fee returns. Spot results also say nothing about funding, borrowing, or liquidation.

The protocol requested prewindow indicator history, but the live MCP schema offered no warmup field and the early snapshot query returned no records. Exact automatic warmup and complete source-bar coverage remain unverified. The saved on-close setting also does not independently establish causal signal-to-fill timing. These runs demonstrate an inspectable workflow under recorded modeling assumptions, rather than validate those assumptions for trading.

Audit the saved rules, including failures

The first actual strategy_create call failed with column StrategyCreationOperation.recoveryOwnerToken does not exist. No successful natural-language conversion is claimed.

The supported alternative exposed by the same live manifest was to clone an owned demo strategy, replace all copied rules through strategy_update, and read the result with strategy_get. The original stayed unchanged. The assistant authored A and B explicitly; strategy_validate_rules returned two valid rules and no errors for each before the updates.

Actual saved candidate A entry showing hourly SMA10 crossing above SMA20, flat-position and pending-buy guards, and fixed 0.01 market buy quantity
Candidate A's saved rule editor after the explicit MCP update. This is assistant-authored rule replacement, not output from the failed natural-language creation request. Open the image to inspect the conditions and quantity.

Cloning produced an ID, not evidence that the intended rules ran. The research record kept each candidate’s validated payload, persisted rule readback, strategy ID, backtest ID, and terminal status separately. Copied baseline statistics were excluded from the comparison.

Compare completed jobs and inspect the loss

Both development jobs reached complete. Their full returned trade lists reconciled to reported P&L within floating-point precision.

Development candidate Trades Reported P&L Reported maximum drawdown
A: SMA10/20 crossover 3 $7.04 0.8567%
B: A plus SMA50 filter 2 −$30.20 0.8599%
Actual completed candidate A report showing $7.04 realized P&L and rounded 0.9 percent maximum drawdown
A's actual report for the January 13–19 development request. The UI rounds the underlying 0.07036% return to 0.1% and 0.8567% drawdown to 0.9%; its score is not a trading recommendation.
Actual completed candidate B report showing negative $30.20 realized P&L, negative 0.30 percent return, and rounded 0.9 percent maximum drawdown
B used the same development dates, sizing, and execution inputs. Adding the filter did not improve this small comparison.

The trade list explained the difference. B excluded A’s $37.23 winner while keeping its $20.19 winner and $50.39 loss. The filter did not reduce reported drawdown here; three trades cannot establish how it behaves generally.

The selected loss entered at $104,539.173465 and exited at $99,500.225, with quantity 0.01 BTC and reported commission zero:

(99,500.225 − 104,539.173465) × 0.01 = −50.38948465
Actual candidate A losing-trade details with entry $104,539.17, exit $99,500.23, account return negative 0.50 percent, and P&L per unit negative $5,038.95
The product displays P&L per unit, approximately −$5,038.95. At 0.01 BTC, the retained trade's total loss is approximately −$50.39; confusing those units would change the conclusion.

The matching native trade record spans January 18 at 19:00 UTC to January 20 at 00:00 UTC, the run boundary. End-of-run flattening was enabled, so this record does not prove the loss closed through the crossover exit. It also prompted a one-day buffer before the reserved evaluation, rather than assuming adjacent date labels could not overlap.

Freeze one refinement, then keep its losing result

A’s largest loss was about 4.82% of the position’s entry value. The assistant retained A as the base and proposed one change: an entry-owned initial stop at order.entry_price * 0.97, anchored to the actual first fill. The original A and B stayed unchanged.

The updated payload was validated and re-read before the reserved January 21–27, 2025 run. That later period had not been inspected during selection. The third job completed with seven losing trades, −$95.70 P&L, and 0.9947% reported maximum drawdown.

Actual completed reserved-period stop-variant report showing negative $95.70 realized P&L, zero percent win rate, and rounded 1.0 percent maximum drawdown
The single frozen refinement's actual later-period result. Different dates prevent treating its P&L difference versus A as proof of improvement or deterioration caused by the stop.

One returned trade lost about 3.0485% of its entry value. A 3% stop trigger therefore did not imply a 3% realized-loss cap in this run. More generally, a triggered stop becomes a market order rather than guaranteeing its trigger price. Investor.gov stop-order explanation

The research stopped after the three prespecified jobs. There was no additional parameter search to replace the losing result. A larger validation program would need untouched data, verified timing and costs, and explicit selection rules; the separate walk-forward example explains that next research stage.

Reuse the evidence standard

Start with a request like this, supplying your own resolved instrument and windows:

Research one hypothesis using my authorized Stratifyre account.
First discover live MCP tools, schemas, resources, and scopes;
verify the intended account and read available limits.
Draft two labeled candidates for the supplied canonical instrument.
Fix dates, warmup, signal/fill timing, capital, sizing, and costs.
Reserve a later period before inspecting results.
Show each complete write/job payload and its consequences for approval.
Use at most one active backtest and three jobs total.
Read saved rules before submitting each job; poll terminal states.
Keep every failure, empty result, loss, strategy ID, and job ID.
Inspect at least one trade and reconcile quantity, prices, fees, and P&L.
Propose one evidence-motivated refinement; freeze it before the reserve.
Separate observed results from unresolved assumptions.
Do not submit live orders, start scanners, or contact recipients.

The reusable output is the candidate-to-result notebook, including rejected ideas and unresolved checks. After this experiment, the owned OAuth tokens were revoked and a request using the revoked access token returned HTTP 401; completed strategies and reports remain available for inspection.

Use the AI assistant connection guides to establish your own supported client, then begin with one bounded question and the same evidence trail.

Put your strategy rules to the test

Build your strategy, inspect historical trades, and review the assumptions behind your results.

Build and test your strategy