We gave Claude Code a job nobody wants to do by hand: a full framework upgrade of a legacy service. The task involved moving a parcel tracking REST API called ShipTrack from Java 11 and Spring Boot 2.7 to Java 25 and Spring Boot 4.1. The initial codebase included legacy elements like Springfox for API docs and the deprecated WebSecurityConfigurerAdapter. The session took about two hours of Claude time and roughly $24 in API fees. While the final build ended green with 47 tests passing, the journey exposed six distinct errors made by the AI. Crucially, none of these mistakes made it into the final commit, thanks to a combination of automated guardrails and human review.

The Illusion of Success in Test Counts

After the move to Java 21, Claude modernized the code and wrote a tidy summary. Everything in it was accurate except one number: it said 56 tests passed. The build had run 47. This discrepancy highlights a fundamental weakness in LLM-based coding agents: the model's summary describes its intent, not necessarily its execution. The team’s solution was to define 'done' not by the model’s self-report, but by a specific command: ./mvnw verify. By forcing the agent to read the actual build output rather than relying on its internal summary, the team caught the hallucinated test count before it could pollute the commit history.

Plan Mode Hallucinations and Metadata Checks

During the migration to Spring Boot 3.5, Claude entered plan mode and generated a detailed roadmap. However, the plan contained two false technical claims: it stated that Spring Boot 3 removed a path-matching property (it didn't) and that server.max-http-header-size retained its name (it was renamed to server.max-http-request-header-size in 3.0). These errors were caught by reviewing the plan like a pull request and verifying against the spring-configuration-metadata.json file shipped within the Spring Boot JARs. This underscores that while plans are useful, they must be validated against authoritative source metadata, not just the model's training data.

Guardrails Forcing Human Decisions

A more subtle failure occurred when two tests failed due to a change in Spring Framework 6.1's static resource handler behavior. Claude correctly identified two options and attempted to pause for human input. However, a custom 'Stop' hook, designed to prevent the agent from quitting on a red build, forced Claude to choose an option itself. The agent picked Option B, got the build green, and flagged the choice. The team later adjusted the hook to allow the agent to stop when only characterization tests are failing, recognizing that behavioral changes in public APIs require human judgment, not just a green build.

Silent API Changes and Configuration Renames

Even with green tests, the public API changed. A side-by-side comparison of the old and new services revealed that while the status code and message matched, the new response body included an extra details field and a different timestamp format. The tests had passed because they only checked for the presence of four fields, not their exclusivity. Later, during the jump to Spring Boot 4.1, Claude misdiagnosed a missing message field in error responses as a Spring Boot change requiring custom code. In reality, the configuration key server.error.include-message had been renamed to spring.web.error.include-message, and Spring Boot was silently ignoring the old key. Both issues were caught by rigorous metadata checks and human comparison of actual HTTP responses.

The Danger of Filtered Output

The final mistake occurred during the Java 25 upgrade. Claude ran a build and checked for warnings using mvn compile | grep -i warning. Finding no matches, it declared success. However, the compilation had actually failed with a cannot find symbol error related to Lombok annotation processing, which had changed in JDK 23. The pipe command swallowed the exit code, and the error message didn't contain the word 'warning'. A subsequent clean build revealed the failure. This incident reinforces a critical lesson for AI-assisted development: filtered output is not a result. The team updated their hooks to force a clean build whenever pom.xml changes, ensuring no stale artifacts mask compilation errors.

Key Takeaways

  • Self-Reporting is Unreliable: Models often hallucinate test counts or success states; always validate against raw build output.
  • Metadata is King: Spring Boot configuration changes must be verified against spring-configuration-metadata.json, not just model knowledge.
  • Green Builds Lie: Tests only protect what they assert; silent API changes can slip through if assertions are not strict.
  • Guardrails Need Escape Hatches: Automated hooks that force a green build can override necessary human decisions on API behavior.
  • Clean Builds are Mandatory: Filtered commands can hide critical compilation errors; always require a clean build for dependency changes.

The Bottom Line

Claude Code did almost all of the work, and did it well, but the human partβ€”building the guardrails and making the final decisionsβ€”was the part that mattered. No single guardrail caught everything; it was the combination of automated checks and human review that kept the six mistakes out of the commit history.