The era of passive meeting AI is ending. Vibeconferencing, an open-source project by Stan James, transforms Google Meet from a broadcast channel into a collaborative workspace for autonomous agents. Unlike traditional tools that generate summaries after the fact, this Electron-based application allows models like Claude Code, Codex, Grokbot, and Meta's Muse to join calls as active participants. The agent listens, speaks back using synthesized voice, and can execute tasksβsuch as drafting emails, writing code, or updating a shared whiteboardβwhile the meeting is still in progress.
Architecture and Real-Time Execution
The technical implementation bypasses direct WebRTC handling by the LLM. Instead, the agent communicates with a bundled MCP server, which then interfaces with the Electron app to manage audio input, virtual camera output, and UI automation within Google Meet or Slack. This separation of concerns allows the agent to focus on reasoning and generation, while the local app handles the media plumbing. The result is a bot that can literally 'show its work' by sharing its screen, rendering avatars, or displaying live code changes to other participants.
From Transcript to Action
The core value proposition is immediacy. In a standard workflow, users copy-paste transcripts into Claude mid-call to get answers. Vibeconferencing eliminates this friction by keeping the context window open and the agent present. Users can issue natural language commands like 'Put a summary of what we decided on the whiteboard' or 'Take notes with diagrams,' and the agent responds instantly. The project supports multiple agents in a single room; recent demos showcased six agents from four different model families interacting simultaneously, though turn-taking remains a complex challenge.
Known Limitations and Experimental Status
As an early-stage experimental tool, Vibeconferencing has notable rough edges. The default voice output is robotic, requiring users to integrate ElevenLabs or local engines like Kokoro for natural speech. Turn-taking between multiple bots can lead to cut-off replies, and the shared whiteboard feature currently risks overwrite conflicts if two agents write simultaneously. Additionally, connection stability is variable, with some joins dropping seconds after admission in certain Workspace-hosted rooms. The project is MIT-licensed and available for macOS, Linux, and Windows, but users should expect a DIY experience.
Key Takeaways
- Vibeconferencing enables AI agents to join Google Meet and Slack as active, speaking participants rather than passive note-takers.
- The app uses an MCP server to bridge LLM reasoning with local media handling, allowing agents to execute tasks and share screens in real-time.
- Supports multiple model families including Claude Code, Codex, Grok, and Meta's Muse, though multi-agent turn-taking remains unstable.
- Requires external API keys for high-quality voice synthesis and faces known connection issues in some Google Workspace environments.
The Bottom Line
Vibeconferencing represents a significant leap from 'meeting notes' to 'meeting participants,' but its experimental nature and reliance on external APIs for quality voice make it a tool for early adopters willing to debug their own digital colleagues.