Tommy Tran on What It Takes to Make AI Agents Work in the Real World
Specializing in generative AI and AI-agent systems, his work focuses on three practical factors that determine whether an agent can be trusted with real work: reliability, cost, and control.
At Ai4 2026, Tommy Tran walked an audience of artificial intelligence (AI) practitioners through a question that polished agent demonstrations can obscure: How should an autonomous system move from a plausible proposal to a production change that engineers can trust?
In the workflow he presented, an AI agent gathers evidence and proposes a change, but the proposal is only the beginning. An engineer reviews it, and a limited rollout measures it against a predefined baseline and safety guardrails. A validated improvement can advance, while an inconclusive or harmful change is held for review or rolled back.
As a software engineer at Meta, that emphasis on restraint runs throughout Tran’s work. He focuses on what must surround a model before a company can depend on it. In his approach, the objective is not maximum autonomy. It is useful autonomy operating inside defined boundaries.
The question is increasingly relevant as businesses give agents access to operational tools and data. A 2026 Deloitte survey of 3,235 technology and business leaders across 24 countries found that 74% expected their companies to be using AI agents at least moderately by 2027. Only 21% said their organizations had mature governance for agentic AI.
Tran’s approach begins with three principles. Define what the agent may do before it starts. Measure the outcome instead of trusting a plausible response. When the evidence is weak or an action could be difficult to reverse, require the agent to stop or turn the decision over to a person.
Human review, staged rollouts, and rollback are established engineering practices. Tran brings them together in a framework organized around an agent’s authority, the evidence it must produce, and the conditions under which it must stop.
Encoding expertise at hyperscale
Tran’s background spans high-growth product engineering and hyperscale infrastructure. At financial technology company Ramp, he built machine-learning systems and data pipelines supporting sales and revenue operations. His more recent work focuses on AI compute and capacity efficiency at hyperscale.
Although the applications differ, both environments place an emphasis on operational outcomes. A model can perform well in isolation while still being too expensive, unreliable, or poorly integrated to improve the larger process.
That distinction becomes especially visible in infrastructure engineering. Automated tools can continuously identify inefficient code or wasted computing resources. The bottleneck comes next: an experienced engineer must investigate the problem, reconstruct the relevant context, decide what can safely change, and determine how to measure whether the change helped.
In the architecture Tran presented at Ai4, agents gather profiling data, search code, retrieve documentation, and examine examples of previous work. They then apply versioned instructions for a particular type of optimization and produce a candidate change with its supporting evidence.
The system is designed to assist with expert judgment rather than conceal or eliminate it. An engineer still reviews the proposal, and a limited rollout measures the result against a predefined baseline and safety guardrails before the change can advance.
An Engineering at Meta article co-authored by Tran reports that automating this kind of diagnosis can compress approximately 10 hours of manual investigation into approximately 30 minutes.
The larger change is not simply speed. The agent handles repetitive context gathering, while the engineer begins with a bounded proposal and the evidence supporting it. Human attention remains concentrated on decisions where judgment carries the most value.
InfoQ later described the platform Tran developed as a significant step toward self-optimizing infrastructure at hyperscale, highlighting its use of reusable agent skills to capture and scale engineering knowledge.
The underlying principle can extend beyond infrastructure: identify the repeatable portion of an expert’s judgment, give an agent limited authority over that portion, and preserve human control over decisions carrying significant consequences.
Making waste visible
Tran also develops open-source projects independently. One is TraceBurn, a local-first profiler for AI agents that records model calls and reports token use, latency, and estimated cost by trace.
The tool allows an engineer to identify where an agent consumed resources and compare its behavior before and after a change. Its purpose is not to replace a company’s complete observability platform. It addresses a narrower question: When an agent costs more than expected, can an engineer locate the waste and verify that a proposed fix worked?
TraceBurn reflects Tran’s view that reliability and efficiency should be measured rather than assumed. A system that generates a convincing response has not necessarily completed its task efficiently. It may have repeated calls, carried unnecessary context, entered a retry loop, or used a more expensive model than the task required.
Before changing the system, an engineer needs to see what happened. After changing it, the engineer needs a meaningful comparison. Without both, an apparent improvement may be little more than an impression.
The same reasoning applies outside software engineering. An agent assisting a customer-service team might gather order history, check company policy, and prepare a response while escalating high-value or ambiguous cases to a person. The company could then measure resolution time, incorrect actions, and escalation rates rather than judging the agent by how polished its messages sound.
For Tran, the opportunity is often not to replace an expert completely. It is to remove repetitive work from the expert’s path while keeping consequential decisions under appropriate human control.
Designing for failure
A second open-source project makes Tran’s thinking about failure more explicit.
His Agent Fault-Tolerance Harness runs a small tool-using agent against a fixed set of tasks while deliberately injecting timeouts, malformed responses, and ambiguous failures. It then compares what happens when established reliability controls are present or absent.
Those controls include deadlines, validation, idempotency, fallback paths, and compensating actions. Each addresses a different part of the failure process, and one safeguard cannot necessarily replace another.
A retry, for example, may recover a failed request. But if the remote action actually succeeded before the connection failed, repeating it may create a duplicate. An idempotency safeguard can prevent that duplicate, while a compensating action may be needed to clean up partial state created elsewhere.
Improving a prompt may not fully address this type of problem; protections may also need to be built into the software and systems surrounding the model.
The harness is deliberately synthetic and is not presented as a universal benchmark. Its controlled setting is what makes it useful. By holding the tasks constant and injecting defined faults, the experiment examines how an agent behaves when its dependencies stop cooperating.
This approach treats deadlines, retries, validation, rollback, and human escalation as part of the product rather than cleanup work. It also places a practical limit on the word autonomous. An agent may be capable of planning and calling tools while still operating inside a carefully constrained system.
The distinction becomes more important as agents gain permission to act rather than merely provide information. An inaccurate paragraph can be corrected. An incorrect transaction, deleted record, or duplicated action may be considerably harder to reverse.
A consistent thesis
Across his infrastructure work and independent tools, Tran returns to three questions that a team should answer before deploying an agent:
- What exactly may the agent change?
- What observable result will count as success?
- What condition will force review, rollback, or shutdown?
If a team cannot answer those questions, its system may be ready for a demonstration but not for consequential work.
The framework challenges the assumption that every human checkpoint is an inefficiency AI should eventually eliminate. Some checkpoints exist because a decision is unusual, ambiguous, or costly to reverse.
In those situations, the most useful role for an agent may be to prepare the decision so thoroughly that an experienced person can make it quickly and with better evidence.
Tran presented this approach at Ai4 2026 and is scheduled to present the infrastructure work at the WeAreDevelopers World Congress in San José in September.
The central idea remains consistent across those presentations and his public projects. As AI moves from experiment to infrastructure, model capability is only part of the equation. Reliability, cost, and control increasingly determine whether an agent can be trusted with real work.
In Tran’s approach, trust does not come from a model’s apparent confidence. It comes from the boundaries around the model, the evidence the system preserves, and its ability to stop.
Entrepreneur Media was not involved in the creation of this content.