<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Darien Williams]]></title><description><![CDATA[Darien Williams]]></description><link>https://darien-ai.hashnode.dev</link><image><url>https://cdn.hashnode.com/res/hashnode/image/upload/v1593680282896/kNC7E8IR4.png</url><title>Darien Williams</title><link>https://darien-ai.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Sun, 11 Oct 2026 06:41:51 GMT</lastBuildDate><atom:link href="https://darien-ai.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Building Reliable AI Agents: From RAG to Tool Calling in Production]]></title><description><![CDATA[Building an AI agent that works in a demo is relatively easy. Building one that I would trust inside a production application is a very different engineering problem.
In my experience working with LLM]]></description><link>https://darien-ai.hashnode.dev/building-reliable-ai-agents-from-rag-to-tool-calling-in-production</link><guid isPermaLink="true">https://darien-ai.hashnode.dev/building-reliable-ai-agents-from-rag-to-tool-calling-in-production</guid><category><![CDATA[AI]]></category><category><![CDATA[llm]]></category><category><![CDATA[agentic AI]]></category><category><![CDATA[RAG ]]></category><dc:creator><![CDATA[Darien Williams]]></dc:creator><pubDate>Thu, 01 Oct 2026 18:03:26 GMT</pubDate><content:encoded><![CDATA[<p>Building an AI agent that works in a demo is relatively easy. Building one that I would trust inside a production application is a very different engineering problem.</p>
<p>In my experience working with LLM applications, RAG, tool calling, and agentic workflows, one of the biggest lessons has been that the model itself is only one part of the system. Reliability comes from the architecture around it.</p>
<p>Start With the Workflow, Not the Model</p>
<p>Before choosing a model or agent framework, I try to understand what the user is actually trying to accomplish.</p>
<p>For example, imagine an assistant that needs to answer questions using internal business information. A simple implementation might send the user's question directly to an LLM.</p>
<p>That works until the model needs information it doesn't have.</p>
<p>A more useful workflow might look like:</p>
<p>User Request → Intent/Planning → Retrieval → Tool Selection → Tool Execution → Validation → Response</p>
<p>Breaking the workflow into steps makes the system easier to understand, test, and improve.</p>
<p>Ground the Agent With RAG</p>
<p>One problem with LLMs is that they can generate convincing answers even when they don't have enough information.</p>
<p>Retrieval-Augmented Generation helps address this.</p>
<p>Instead of expecting the model to know company-specific information, I retrieve relevant information from trusted sources first.</p>
<p>A typical RAG pipeline might include:</p>
<p>Receive the user's question.</p>
<p>Create an embedding for the query.</p>
<p>Search a vector database for relevant information.</p>
<p>Add the retrieved context to the model request.</p>
<p>Ask the model to answer using that context.</p>
<p>Technologies such as pgvector, Pinecone, Weaviate, or Milvus can support the retrieval layer.</p>
<p>The important part isn't simply adding a vector database. Retrieval quality needs to be evaluated just like model quality.</p>
<p>Give Agents Tools, Not Unlimited Access</p>
<p>Agents become much more useful when they can interact with real systems.</p>
<p>Instead of only generating text, an agent might be able to:</p>
<p>Search internal knowledge.</p>
<p>Query a database.</p>
<p>Call an internal REST API.</p>
<p>Retrieve account information.</p>
<p>Trigger an approved workflow.</p>
<p>Generate structured data for another service.</p>
<p>I prefer exposing small, clearly defined tools rather than giving an agent broad access to a system.</p>
<p>For example:</p>
<p>get_account_information(account_id)</p>
<p>is much easier to validate and secure than allowing an agent to construct arbitrary database queries.</p>
<p>Each tool should have a clear input schema, output schema, permissions, timeout behavior, and error-handling strategy.</p>
<p>Use Structured Outputs</p>
<p>Free-form model output can create problems when another service needs to consume the result.</p>
<p>Whenever possible, I prefer structured responses.</p>
<p>Instead of:</p>
<p>"Looks like we should retrieve the customer's transactions."</p>
<p>the model can return something conceptually like:</p>
<p>{ "action": "get_transactions", "account_id": "12345", "date_range": "30_days" }</p>
<p>The application can validate that output before executing anything.</p>
<p>This creates a clean boundary between probabilistic model behavior and deterministic application logic.</p>
<p>Keep Business Logic Outside the LLM</p>
<p>Another lesson I've learned is not to put every decision inside a prompt.</p>
<p>Important business rules should remain deterministic application code whenever possible.</p>
<p>The model can help interpret an ambiguous user request, identify relevant information, or select an appropriate tool.</p>
<p>But authorization, validation, financial calculations, permissions, and other critical rules should generally remain outside the model.</p>
<p>This makes the system easier to test and reduces unexpected behavior.</p>
<p>Evaluation Is Part of the Product</p>
<p>Traditional software can often be tested with simple expected outputs.</p>
<p>LLM systems require another layer of evaluation.</p>
<p>I like to maintain representative test cases covering things such as:</p>
<p>Correct tool selection.</p>
<p>Retrieval relevance.</p>
<p>Grounded responses.</p>
<p>Structured-output validity.</p>
<p>Hallucination behavior.</p>
<p>Failure scenarios.</p>
<p>Latency.</p>
<p>Token usage.</p>
<p>When prompts, models, retrieval strategies, or tools change, these cases can be rerun to detect regressions.</p>
<p>Observability Matters</p>
<p>When an agent fails, "the AI gave a bad answer" isn't enough information.</p>
<p>A production system should make it possible to understand what happened.</p>
<p>Useful information includes:</p>
<p>User request.</p>
<p>Retrieved context.</p>
<p>Model selected.</p>
<p>Tool calls.</p>
<p>Tool results.</p>
<p>Model latency.</p>
<p>Token usage.</p>
<p>Validation failures.</p>
<p>Final outcome.</p>
<p>This makes debugging agent behavior much closer to debugging a distributed application.</p>
<p>Design for Failure</p>
<p>External APIs fail.</p>
<p>Models time out.</p>
<p>Retrieval sometimes returns poor context.</p>
<p>Tools can return unexpected responses.</p>
<p>Production AI systems need to expect these situations.</p>
<p>Depending on the workflow, I use techniques such as timeouts, retries, validation, fallback behavior, idempotency, and clear error states.</p>
<p>An agent should fail predictably rather than continuing to make increasingly uncertain decisions.</p>
<p>Final Thoughts</p>
<p>The most interesting part of AI engineering for me isn't simply connecting an application to an LLM.</p>
<p>It's designing the system around the model.</p>
<p>RAG, tool calling, structured outputs, evaluation, observability, and deterministic business logic turn a model into something that can become part of a reliable software product.</p>
<p>My approach is simple: let the model handle the problems where probabilistic reasoning is useful, and let traditional software engineering provide the boundaries that make that reasoning safe and dependable.</p>
<p>That combination is what makes production AI systems interesting to build.</p>
]]></content:encoded></item></channel></rss>