Agent Bricks, Revisited: How Agents Are Built on Databricks Today

Updated Date:

When I wrote about Agent Bricks in June 2025, the pitch was simple. Describe the task, let Databricks handle evaluation and optimization, and improve quality through human feedback. Fifteen months later, that idea hasn’t gone away. It now sits inside a much broader agent platform. If you’re evaluating Databricks for agent work today, here’s how I’d describe it.

Start with the models

Everything starts with Foundation Model serving. Databricks hosts a curated set of open-source and third-party models behind secure, scalable APIs, including Meta Llama, Anthropic Claude, and OpenAI GPT. You can pick the model per workload without leaving the governance boundary of your workspace.

Three ways to build

No-code prototyping in AI Playground. Choose an LLM, attach tools, and chat with the agent to see how it behaves. When the behavior looks right, export it to code. This is the fastest way to answer “is this agent idea viable?” before anyone writes Python. Databricks has a no-code getting-started guide that walks through it.

Agent Bricks agent types. This is where the original Agent Bricks vision lives now.

Knowledge Assistant builds a question-and-answer chatbot over your documents and answers with citations. It uses an approach Databricks calls Instructed Retriever, designed to get past the limits of classic RAG (retrieval-augmented generation), and it handles documents in multiple languages. You create one from the Agents page. You then point it at up to 10 knowledge sources: files in a Unity Catalog volume, a Unity Catalog table with a file column, or an AI Search index. It’s served as an endpoint you can call from your applications.

Supervisor Agent handles multi-agent orchestration. It coordinates Genie Agents, other agent endpoints, Unity Catalog functions, MCP servers, and your own custom agents behind one entry point.

Custom agents in Python. For full control, you write the agent yourself using whichever framework your team prefers, such as LangGraph, LangChain, the OpenAI SDK, or LlamaIndex. MLflow Tracing is built in, and Databricks Apps gives you a quick iteration loop; the agent quickstart is the fastest way in. Agent tools can query structured and unstructured data, run code, or call external APIs. MCP (Model Context Protocol) gives agents a standard, secure way to connect to data and tools. If you’re deciding between a single agent and a multi-agent design, the agent system design patterns page is worth reading first.

What became of human feedback

The part of the original announcement I was most excited about was ALHF, the idea that experts could steer an agent in plain language. In practice it now works like this.

In Knowledge Assistant, you add the questions users actually ask, or ones the agent got wrong. You share the configuration page with your SMEs, and they write guidelines for each question (Improve quality). The guidelines apply immediately, so you can retest right away in the builder or in AI Playground. Labeled questions and guidelines can be imported from or exported to Unity Catalog tables, which makes the feedback a reusable, governed asset rather than something trapped in one tool. During development you can also label traces directly in the UI.

For all agents, Agent Evaluation extends the same idea. Stakeholders give feedback through built-in review apps. LLM judges flag quality problems. MLflow measures quality alongside cost and latency. The same judges and custom metrics then run as production monitoring, so what “good” means during development is the same bar you hold in production. That’s a better answer to “auto-optimizing agents” than the original framing: less magic, more measurable. The MLflow 3 for GenAI getting-started guide covers tracing, evaluation, and feedback end to end.

Governance got a real answer

A gap in 2025 was what to do with agents that don’t run on Databricks. Agent services, currently in Beta, let you register an externally hosted agent in Unity Catalog. Teams can discover it from a single view, and access is controlled with the same grants you already use for tables, models, and functions. Agents built inside Databricks use workspace permissions too. Knowledge Assistant, for example, separates Can Manage from Can Query: the first lets someone edit, tune, and set permissions, and the second only lets them call the endpoint.

Practical notes before you build

A few details from the Knowledge Assistant docs that are worth knowing up front:

My take

The original promise of Agent Bricks was that teams could focus on what an agent should do, not on tuning it. That promise mostly held up, but it arrived as a toolkit rather than a single product. Use Knowledge Assistant when the job is grounded Q&A over documents. Use Supervisor Agent when you need to route across Genie, functions, and other agents. Write custom agents when neither fits. Wrap all of it in MLflow evaluation and Unity Catalog governance. The human-feedback loop I highlighted last year is still the differentiator. It’s just called guidelines and review apps now, and it’s backed by judges you can take into production.

References

Leave a comment

Trending