How Deputy runs the AI inside its workforce management platform
Built on Datadog’s Agent Observability platform alongside the monitoring Deputy already ran. The solution captures every conversation, model call, and tool call into a single trace, wires the automated evaluation suite into the same pipeline, and connects the team’s coding agents straight to production data. Deputy now observes, debugs, and cost-manages its assistant with the same platform and largely the same habits it uses for every other service it runs.
- Every answer is traceable to the agent logic that produced it.
- AI spend is monitored in real-time dashboards, not estimated retrospectively.
- Failed evals and incidents are investigated within the same observability platform.
All about Deputy
Deputy, the leading Australian-founded workforce management platform, partnering with Mantel and Datadog, has built out an enterprise-grade operational capability to support their market-leading AI assistant: the tracing, evaluation and debugging tooling that sits underneath the AI that allows their customers to build and publish staff rosters using natural language.
The assistant lets managers work in plain English. It can search and validate shifts, auto-build and publish a roster, approve timesheets and handle timesheet queries, and answer how-to questions by searching Deputy’s help content. Deputy serves 390,000+ customers in over 110 countries.
The engineering behind the scenes
The engineering effort behind it is shared with Mantel, the Australian consultancy whose engineers work embedded in Deputy’s AI team. Much of that joint work has gone into a part of the stack that gets little attention in most AI projects: what happens after the feature ships.
An agent is not a feature that ships and then holds still: people find phrasings nobody wrote down, ask for things in an order the design never anticipated, and change how they ask once they work out what the assistant is good at. None of that is knowable in advance. Researching how the agent actually behaved, conversation by conversation, is what keeps its answers accurate – and what turns up there becomes the next eval.
“Often, the focus is purely on building functionality with AI. However, we are focused on building in scalability of AI solutions from the first agent - mature customers like Deputy are focusing on how they embed observability and traceability into their agents to ensure they can safely and securely scale from a single agent to scaled multi-agent solutions.”
Adam DurbinCTO, Mantel
Tracing every conversation
Each conversation with the assistant is recorded as a trace in Datadog’s Agent Observability product, alongside the APM traces, logs and monitors Deputy already runs on the platform for the rest of its estate.
A single request from a manager fans out into a surprisingly deep tree of steps before an answer comes back, and every one of them lands in the trace: each model call, each search, each tool call against the Deputy platform, along with its inputs and outputs. Model calls are recorded with their token counts, estimated dollar cost, latency, and how much of the prompt was served from cache, which turns spend from something the team estimates into something it can read off a dashboard.
Each span also carries the git commit of the deploy that produced it. Personal information gets handled before a human ever reads a trace. Datadog’s Sensitive Data Scanner, which comes bundled with Agent Observability, picks up things like email addresses in prompts and responses, and what lands in the stored trace is a placeholder naming the rule that matched rather than the address itself, with the span tagged by category. None of that was built by the team. It came with the platform.
For the engineers on the team, the Agent Observability trace explorer has become the default debugging surface. A developer can scan the day’s conversations, see the input and output of each one at a glance, then click into a trace for the full picture: which tools were called, what went into them, what came back, and what each model call produced along the way.
The test suite lands in the same place
On a typical day, a noticeable share of the traces arriving in Datadog are not from users. They come from Deputy’s own evaluation suite.
The team maintains a set of automated evals, single-turn and multi-turn conversation scenarios with the assistant’s tools mocked out and the responses scored. These run in CI, and they emit traces into the same Agent Observability pipeline as production traffic. An engineer investigating a failed eval uses the same views and tooling they would use on a live incident.
The suite is also fed by people. Before releases, staff from across the company put realistic scenarios to the assistant in sessions the team calls test parties, and what they find makes its way back into the automated evals.
“It's the first thing every developer on the team opens. If someone tells us the assistant did something strange, we don't sit around guessing. We pull up the trace and read the whole thing back, step by step. And because the version is stamped on every span, a drop in quality gets chased down the same way we'd chase down any other regression.”
Vihan PatelHead of AI Solutions, Mantel
Letting the coding agents do the digging
The newest operational pattern on the team connects Datadog to the coding agents themselves. Datadog exposes its platform to AI agents through an MCP server and a command line tool called Pup CLI, and Deputy’s engineers have wired both into the coding agent they develop with.
Handed a production bug, the agent searches the relevant logs itself, refines its queries and tries variants when the first pass comes back empty, pulls the offending trace, and proposes a fix from inside the same session where the code gets written and shipped.
“That loop is the part I’d fight hardest to keep,” says Vihan. “The agent writing the fix is the same one reading the production data, so an investigation that used to eat an afternoon of flicking between tabs is basically just a conversation now.”
“For workforce platforms such as Deputy, the value of agentic AI lies in making complex, time-consuming tasks like scheduling and timesheet management simpler for managers and frontline teams. But moving from a successful pilot to AI that can operate at scale requires confidence in its reliability, safety, cost and business impact. Deputy’s work with Mantel shows the value of combining ambitious product thinking with disciplined delivery. Together, Mantel and Datadog provide the practical expertise and unified, AI-powered observability organisations need to scale AI responsibly and turn innovation into dependable customer value.”
Yadi NarayanaField CTO, Datadog
“The Datadog Agent Observability platform has given us complete visibility into agent behavior, cost, and latency. It has allowed us to proactively resolve issues, speed up testing, and deliver a better customer experience. Partnering with Mantel made the implementation smooth and tailored to our specific operational needs.”
Luis SanchezSenior Engineering Manager, Deputy
Consolidating into one platform
Early on, Datadog was paired with a separate, dedicated tool for the long-running record of eval results used to monitor how scores trend across releases. Datadog capability moved quickly since, and the solution evolved. Deputy is now consolidating it as the main platform for the AI product: the whole eval infrastructure is migrating across, and sensitive data redaction comes with it.
Datadog is the backbone for the engineering side, operated by the Mantel Managed Services team. The assistant is observed, debugged and cost-managed with the same platform and largely the same habits the company uses for its ordinary services, with the LLM-specific pieces layered on top rather than built to the side. It’s becoming the central place the conversations themselves are examined, too – monitored for behaviour, tagged and annotated; and not only by engineers. Product managers and subject-matter experts, the people who can tell a good answer from a merely plausible one, review in the same place.