AI agent security is an application design problem as much as a model problem. Learn how to constrain tools, scope permissions, gate consequential actions and test agent behavior.
Quick summary Evaluate the whole graph: assess answers, routing, tool use, reliability and cost. Use deliberate test cases: cover ordinary requests, edge cases and safe failures. Compare against a baseline: combine offline tests with production feedback. Apply focused human review: use a consistent rubric where judgement matters. LangGraph evaluation turns a convincing multi-agent demo into […]
Quick summary Checkpoint state: persistence lets a graph pause, recover and continue. Resume the same thread: reuse its identifier and supply the review decision. Design for replay: repeated execution must not duplicate consequential actions. Use durable storage: production needs an appropriate backend and retention policy. LangGraph persistence is what makes a multi-agent workflow safe to […]
LangGraph routing is the point where a multi-agent graph stops broadcasting work and starts making deliberate choices. It decides which specialist should handle a request, which tasks can run at the same time, and when their findings are ready to combine. In the previous tutorial on LangGraph shared state, we defined safe information flow. Here, […]
Quick summary Define a state contract: make field ownership, inputs and updates explicit. Separate context types: shared facts, private material and long-term memory serve different purposes. Return partial updates: use deliberate reducers instead of mutating a shared snapshot. Test data flow: verify merge behaviour and the context each worker receives. LangGraph shared state is where […]
Quick summary Give coordination a clear owner: a supervisor selects specialists and remains responsible for the final response. Keep workers focused: supply bounded tasks and request concise, structured results. Control the workflow in code: enforce budgets, permissions and stopping conditions outside model instructions. Start simple: introduce specialists only when separate responsibilities make the system easier […]
Quick summary Evaluate the entire trajectory: tool calls and state changes matter as much as an agent’s final answer. Budget for the whole task: repeated retrieval, retries and long-running execution change operational cost. Use layered controls: permissions, checkpoints, monitoring and recovery need deliberate ownership. Keep authority bounded: greater model capability does not replace application-level safeguards […]
If you are comparing MCP vs API, the short answer is that they solve different integration problems. APIs expose capabilities; MCP gives AI applications a standard way to discover and use tools and context. Key takeaway: An API is a contract for interacting with a system. MCP is a protocol that helps AI applications discover […]
Ninety four thousand block reads for one row, by its UUID primary key. I read that line three times, half convinced I had pasted the wrong output into my own terminal. That query used to come back in under 20 milliseconds. At 200 million rows, it took almost four seconds. The only reason anyone noticed […]
An agent evaluation flywheel can expose the path an AI agent took to reach an answer. A final-answer score can hide wrong tools, needless retries, weak evidence, or a failure to adapt when new information appears. That creates a practical question: what should change next? This post turns that diagnosis into a five-stage operating model. […]
Trajectory evaluation shows how a final answer can look perfect while the agent behind it has already failed. Imagine an agent that tells a support team: “The customer record has been updated.” The sentence is clear and reassuring. However, the trace may show a different story. It may select a search tool instead of an […]
If by “turbo vec” you mean the project source, it is a project aimed at a familiar AI-infrastructure problem: storing and searching embedding vectors without treating memory, storage, and retrieval quality as afterthoughts. Its repository describes TurboVec as a vector index built on TurboQuant, written in Rust, with Python bindings. That description is useful, but […]
This multi-agent workflow roadmap introduces coordinated AI roles and workflows working toward one outcome—not magical, fully autonomous AI teams. It is the starting point for the series, and it will become a linked learning path as each tutorial is published. Key takeaways A multi-agent system divides a broader job among defined AI roles and workflow […]
A feature request looks small until it crosses every layer of an application. Clean Architecture with SOLID helps when a checkout rule touches an HTTP handler, ORM model, pricing calculation, email notification and message consumer at once. It separates the business decision from the machinery used to deliver and store it. Clean Architecture is not […]