Developers often reflexively send flaky test failures to remote AI models or CI servers, but a new tutorial argues this is a mistake. The proposed solution is a local Python script that runs a test fixture twice to determine if the failure class is stable. If the error type changes between the two local runs, the script outputs a 'hold' action, preventing any network calls or server boots.
The Two-Run Stability Check
The core mechanism involves executing a specific test script twice using subprocess.run with a 15-second timeout. The script captures the stderr, extracts the exception type using regex, and maps it to one of four predefined classes: local_typo, contract, needs_runtime, or unknown. If the exception class differs between the first and second run, the decision logic returns hold. This ensures that non-deterministic failures are caught before they consume resources on a remote machine.
Why Free Isn't Always Harmless
The author emphasizes that 'free' model access or server usage is not synonymous with 'safe' or 'harmless.' Sending unstable code to a remote service can leak paths, tokens, or secrets into third-party logs. The tutorial explicitly states that if the two local runs do not share the same failure class, you should not boot a box or send the fixture anywhere. This discipline forces developers to fix local_typo or contract errors locally, where the context is fully controlled.
Implementation Details and Limitations
The script writes its decision to ledger/latest.json, providing a diffable artifact rather than a transient chat log. The author notes significant limitations: the approach assumes a Unix-like environment for path handling, does not account for wrapped exceptions that might mask the true error type, and has not been tested on Windows. Additionally, the tutorial is sponsored by MonkeyCode, but the author deliberately avoids citing specific model IDs or token limits, urging readers to check current product docs instead.
Key Takeaways
- Stability across two local runs is the primary gate for remote intervention.
- Four error classes (
local_typo,contract,needs_runtime,unknown) dictate workflow actions. - Remote retries are only eligible when the
needs_runtimeclass is stable across runs. - The script prevents data leakage by ensuring only stable, isolated fixtures leave the local machine.
The Bottom Line
Stop treating free compute as a default answer to every red test. If your failure class flips between runs, you do not have a bug yet; you have a mystery. Hold the retry, fix the instability, and only send stable fixtures to the cloud.