A developer building a customer support agent on Amazon Bedrock AgentCore discovered that high LLM-as-a-judge scores are dangerously misleading for agentic workflows. Despite achieving a consistent 0.92 correctness rating across two independent evaluation runs, the chatbot frequently hallucinated ticket IDs and wrote junk data to DynamoDB, behaviors completely invisible to text-based scoring metrics.

The Architecture and the Illusion of Success

The project, part of the Udacity ร— AWS Agent Engineer Nanodegree, utilizes Amazon Nova Pro with greedy decoding (temperature 0, topK 1) to route messages via a single system prompt. The architecture replaces traditional classifiers with a prompt-driven routing system that feeds into an AgentCore managed harness, which executes tools via an MCP gateway connected to Lambda and DynamoDB. The developer, Mialy, implemented 15 explicit tie-breakers to handle ambiguous inputs, such as distinguishing between a declined card (platform question) and a checkout crash (bug report). The evaluation suite, comprising 13 single-turn cases, relied on Nova Pro to score responses against reference answers, yielding a perfect 1.00 on 12 cases and 0.00 on one, averaging to that impressive 0.92.

Side Effects Are the Real Ground Truth

The critical failure mode was not in the chat text, but in the database state. In a multi-turn conversation about a broken search bar, the bot confidently told the customer, "I have filed a bug report with ticket ID TIX-345678." However, the DynamoDB table remained unchanged, and no tool call appeared in the terminal logs. The model had invented a placeholder-shaped string rather than returning the actual UUID generated by the Lambda function. This "invented ticket ID" defect was only caught by manually comparing the chat transcript against the database rows, a step the automated evaluation pipeline completely skipped.

Prompt Engineering Limitations and Fixes

Another defect involved the bot saving its own clarifying questions as customer-provided data. When a user said, "The app keeps logging me out," the bot filled the stepsToReproduce field with filler text like "Using the app" or its own question, "Please provide specific actions...". The prompt's leniency in accepting short descriptions caused the model to treat the fault description as the reproduction steps when they were semantically identical. The developer noted that while the prompt instructed the model to "never fabricate information," this was insufficient. The fix required explicit rules: the only valid ticket ID is the exact string returned by a successful tool call, and assistant-generated text must never be stored in customer fields.

Key Takeaways

  • Assert on side effects: For agents that write to databases or APIs, automated tests must verify the stored row, not just the chat response.
  • Single-turn evals are insufficient: The 13-case suite used single-turn prompts, but the most critical bugs emerged in multi-turn interactions where state accumulation mattered.
  • Prompt position matters: Rules buried in the middle of long prompts were inconsistently followed even with greedy decoding; moving critical constraints to the top improved adherence.
  • AgentCore gotchas: Gateway target names must avoid dashes to prevent invalid tool sequences, and session IDs require a minimum of 33 characters.

The Bottom Line

If your agent touches a database, your evaluation suite is broken unless it checks the database. Stop trusting text scores and start asserting state changes.

Practical Implementation Details

The developer highlighted several AgentCore-specific pitfalls that consumed debugging time. For instance, the update_harness API requires a different memory configuration shape than create_harness, a subtle discrepancy that caused script failures on re-runs. Additionally, the Gateway passes tool arguments directly as the Lambda event without the traditional Agents Classic parameters envelope, requiring developers to extract the tool name from context.client_context.custom. These infrastructure nuances, combined with the model's tendency to leak reasoning tokens like , demonstrate that building reliable agents is as much about infrastructure plumbing as it is about prompt engineering.