Kevin Meneses Gonzalez recently subjected an AI-generated investment strategy to rigorous scrutiny, moving beyond the standard single-date backtest to a rolling-window analysis. The experiment tested a $10,000 portfolio allocated by Claude across 105 different start dates from January 2015 to September 2023. The methodology utilized adjusted close prices from the EODHD API to account for dividends and splits, ensuring data integrity across the twelve tickers involved. This approach transforms anecdotal performance into statistical evidence, challenging the validity of static backtests commonly found in financial commentary.
The Dominance of Timing Over Selection
The results highlight a stark disparity between the impact of asset selection and market entry timing. Across the 105 three-year windows, the median difference between Claude's portfolio and the S&P 500 benchmark was merely $637. In contrast, the variance between the best and worst start dates for the same portfolio reached $6,753. This tenfold difference demonstrates that the calendar dictates outcomes far more aggressively than the specific mix of assets. For developers building financial agents, this suggests that robust backtesting tools must prioritize distributional analysis over point-in-time snapshots to provide meaningful insights.
Performance Metrics and Drawdown Analysis
Claude's portfolio, consisting of 35% VOO, 20% MSFT, 20% JNJ, 15% GLD, and 10% SHY, outperformed the S&P 500 in only 54% of the tested start dates. However, the true advantage lay in risk mitigation rather than raw return generation. The median maximum drawdown for the AI-allocated portfolio was -17.7%, significantly lower than the -24.5% observed for the S&P 500 and the -21.7% for a traditional 60/40 fund. This indicates that the inclusion of gold and short-term treasuries served to cushion downside volatility during poor entry periods, effectively reducing the psychological burden of holding through market declines.
The Hindsight Bias of Concentrated Bets
The analysis also examined the Magnificent 7 portfolio, which delivered a median ending value of $25,574 and beat the S&P 500 in 100% of start dates. Despite this perfect win rate, the strategy suffers from severe look-ahead bias, as the group was selected based on post-2015 performance. The volatility of this concentrated bet was extreme, with a range of $43,339 between the best and worst outcomes depending solely on the entry month. Furthermore, the portfolio experienced a worst-case drawdown of -50.7%, illustrating that high returns often come with risks that most retail investors are ill-equipped to handle.
Dollar-Cost Averaging vs. Lump Sum Investing
To address timing risk, the study compared lump-sum investing against a six-month dollar-cost averaging (DCA) strategy. Contrary to popular belief, DCA only outperformed the lump-sum approach in 22% of the start dates. While DCA provided a modest buffer during peak entries like February 2020, it consistently underperformed in rising markets. The data reinforces the conventional wisdom that markets tend to rise over time, making immediate capital deployment statistically superior for return optimization, though DCA remains a valid tool for managing investor regret.
Key Takeaways
- Entry timing impacts portfolio outcomes roughly ten times more than asset allocation choices.
- Diversified portfolios with alternative assets reduce median drawdowns by approximately one-third compared to pure equity benchmarks.
- The Magnificent 7's 100% win rate is an artifact of hindsight bias and does not guarantee future outperformance.
- Lump-sum investing outperforms dollar-cost averaging in 78% of historical three-year windows.
The Bottom Line
Stop obsessing over which AI picks the best stocks; focus on building tools that help investors withstand the inevitable volatility of their entry date. The portfolio is secondary to the timing luck, and a robust agent must account for the distribution of outcomes, not just the average.