A developer has released vmtest, a utility designed to force Claude Code to execute UI testing within a Hyper-V virtual machine rather than on the host PC. This addresses a persistent friction point for users relying on screen readers like JAWS or NVDA, where default AI agent behavior hijacks keyboard focus and interrupts workflow with unexpected window changes.

The Problem With Host-Side Testing

When Claude tests Windows desktop applications directly on the user's machine, it inevitably causes UI collisions. Pop-ups steal focus, keystrokes intended for the test app land in active user windows, and screen readers vocalize test artifacts mid-sentence. The author describes this as "chaos," noting that the constant interruptions made it impossible to continue working while the AI agent performed its verification loops.

How vmtest Works

vmtest leverages PowerShell Direct, a Hyper-V feature that requires no network configuration or visible windows. A lightweight helper inside the VM launches the target application, inspects its accessibility tree, and simulates user interactions such as button presses and keystrokes. After each step, the helper reports the current keyboard focus state back to Claude, allowing the model to verify UI logic without touching the host’s input devices.

Integration and Workflow

The tool integrates with Claude Code via a skill file that teaches the model to automatically route window-opening tests to the VM. Users can enforce this behavior globally by adding a directive to their CLAUDE.md file, ensuring Claude never runs intrusive tests on the host. The system uses checkpoints to maintain state per repository branch, allowing the agent to pause and resume testing with the application still open, while merging code returns the VM to a clean state.

Key Takeaways

  • vmtest requires Windows 11 Pro, Enterprise, or Education due to Hyper-V dependencies; Home edition is unsupported.
  • The solution uses a throwaway password for the VM, emphasizing security isolation for testing environments.
  • Testing via VM provides faster feedback than CI pipelines while avoiding the focus-stealing issues of host testing.
  • The author successfully used two Claude sessions to debug vmtest itself, demonstrating recursive AI-assisted development.

The Bottom Line

This is a pragmatic fix for the "last mile" problem in AI coding agents. While LLMs are excellent at logic, their physical interaction with GUIs remains disruptive; vmtest proves that sandboxing is the necessary bridge between autonomous testing and human productivity.

Limitations and Requirements

The current implementation is strictly for Windows desktop apps and requires significant memory resources, consuming between 2 to 8 GB while active. The author notes that while vmtest catches many issues, it cannot fully replace human verification with diverse assistive technologies. Users must manually enable Hyper-V and add themselves to the Hyper-V Administrators group to avoid constant elevation prompts during operation.