finds.dev← search

// the find

the-ai-merge/multimodal-agents-course

★ 581 · Python · Apache-2.0 · updated Jan 2026

An MCP Multimodal AI Agent with eyes and ears!

A five-module course (not a library) that walks through building an MCP-based multimodal agent for video: a Pixeltable-backed MCP server for video/image/audio processing, a Groq-powered tool-use agent with a hand-rolled MCP client, and a React chat UI on top. Aimed at ML/software engineers who want a full worked example of MCP tool-calling beyond wiring an existing server into Claude Desktop.

Splits into three real services (kubrick-mcp, kubrick-api, kubrick-ui) with docker-compose for dev and prod, so it reads like an actual deployable system instead of a notebook. It implements its own MCP client and tool-to-provider translation layer (mapping MCP tools to Groq's Llama 4 tool format) rather than just consuming someone else's server. Opik tracing and prompt versioning are wired in end-to-end, which most agent tutorials skip. FastMCP usage covers resources, prompts, and tools, not just the tools half of the protocol.

It's a course, and the code leans on the linked Substack articles for the actual explanations — reading the repo alone leaves gaps in why things are structured this way. Model choice is hard-locked to Groq (Llama 4) and OpenAI, so swapping providers means rewriting the tool-translation layer, not just a config change. Getting it running requires a separate GETTING_STARTED.md and several API keys (Groq, OpenAI, Opik), so it's not a quick clone-and-run despite the 'minimum cost' framing. Being a two-author course project tied to a specific video series, there's no signal of ongoing maintenance once the syllabus is finished.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →