# Amigo AI > Build, train, and deploy clinical AI agents for healthcare. Amigo supports patient engagement, clinical workflows, and care between visits. Amigo provides a platform for healthcare organizations to configure clinical agents, train them through Digital Residency, integrate with clinical systems, and monitor their performance. The linked pages are the primary sources for product details and customer outcomes. ## Company and platform - [Home](https://www.amigo.ai) - [Clinical AI agent platform](https://www.amigo.ai/platform) - [About Amigo AI](https://www.amigo.ai/about) - [Customers](https://www.amigo.ai/customers) ## Clinical workflows - [Scheduling And Reminders](https://www.amigo.ai/solutions/scheduling-and-reminders) - [Intake And Triage](https://www.amigo.ai/solutions/intake-and-triage) - [Pre Visit Preparation](https://www.amigo.ai/solutions/pre-visit-preparation) - [Pre Visit Summary](https://www.amigo.ai/solutions/pre-visit-summary) - [Clinical Scribe And Copilot](https://www.amigo.ai/solutions/clinical-scribe-and-copilot) - [Care Plan Generation](https://www.amigo.ai/solutions/care-plan-generation) - [Post Visit Follow Up](https://www.amigo.ai/solutions/post-visit-follow-up) - [Between Visit Support](https://www.amigo.ai/solutions/between-visit-support) - [Medication Management](https://www.amigo.ai/solutions/medication-management) - [Proactive Outreach](https://www.amigo.ai/solutions/proactive-outreach) - [Referral Management](https://www.amigo.ai/solutions/referral-management) ## Healthcare specialties - [Primary Care](https://www.amigo.ai/specialties/primary-care) - [Behavioral Health](https://www.amigo.ai/specialties/behavioral-health) - [Cancer Care](https://www.amigo.ai/specialties/cancer-care) - [Digital Health](https://www.amigo.ai/specialties/digital-health) - [Orthopedic](https://www.amigo.ai/specialties/orthopedic) - [Wellness and longevity](https://www.amigo.ai/specialties/wellness-longevity) - [Women's health](https://www.amigo.ai/specialties/womens-health) ## Customer stories and partnerships - [Databricks](https://www.amigo.ai/partners/databricks) - [Eucalyptus](https://www.amigo.ai/customers/eucalyptus) - [CaringHand](https://www.amigo.ai/customers/caringhand) - [Jasper Health](https://www.amigo.ai/customers/jasper-health) - [Nortal](https://www.amigo.ai/customers/nortal) - [The Care Clinic](https://www.amigo.ai/customers/the-care-clinic) - [Nortal partnership (Arabic)](https://www.amigo.ai/arabic/customers/nortal) ## Guides and resources - [Blog](https://www.amigo.ai/blog) - [Guides](https://www.amigo.ai/guides) - [First Use Case](https://www.amigo.ai/guides/first-use-case) - [Healthcare AI ROI guide](https://www.amigo.ai/guides/healthcare-ai-roi) - [AI vendor evaluation guide](https://www.amigo.ai/guides/ai-vendor-evaluation) ## Published articles - [Why I'm Joining Amigo as Chief Medical Officer](https://www.amigo.ai/blog/dr-jay-shah-amigo-chief-medical-officer): I became a cancer surgeon to help people through their hardest days. I'm joining Amigo so I can reach more people and multiply my impact. Published 2026-06-08; author: Jay Shah, MD. [Markdown](https://www.amigo.ai/blog/dr-jay-shah-amigo-chief-medical-officer/markdown) - [Too Many Apps, Not Enough Medicine: The Platform Approach](https://www.amigo.ai/blog/amigo-platform-approach): Why healthcare's digital fragmentation is a patient safety crisis, and what a platform can do about it Published 2026-05-08; author: Ross Green, MD. [Markdown](https://www.amigo.ai/blog/amigo-platform-approach/markdown) - [How We Build Self-Improving Agents](https://www.amigo.ai/blog/agent-forge): Our answer to the post-launch optimization problem Published 2026-04-16; author: Claire Uhm. [Markdown](https://www.amigo.ai/blog/agent-forge/markdown) - [Amigo AI Raises $11M Series A to Train Clinical AI Agents Like Doctors](https://www.amigo.ai/blog/amigo-series-a): This brings Amigo's total funding raised to $17M. Published 2026-03-10; author: Ali Khokhar. [Markdown](https://www.amigo.ai/blog/amigo-series-a/markdown) - [How Clinical Agents Get Built and Trained](https://www.amigo.ai/blog/how-clinical-agents-get-built-and-trained): A walk through of what agents are and how Amigo builds them Published 2026-02-27; author: Nikhil Krishnan. [Markdown](https://www.amigo.ai/blog/how-clinical-agents-get-built-and-trained/markdown) - [CMS ACCESS Model Explained (And How Amigo Helps You Qualify)](https://www.amigo.ai/blog/cms-access-model-explained): What would your chronic care program look like if Medicare paid you to modernize it? Published 2026-02-13; author: Claire Uhm. [Markdown](https://www.amigo.ai/blog/cms-access-model-explained/markdown) - [The Third Leading Cause of Death is Preventable](https://www.amigo.ai/blog/actions-architecture): How Amigo's Actions architecture enables AI agents to safely and reliably execute clinical tasks. Published 2026-01-05; author: Claire Uhm. [Markdown](https://www.amigo.ai/blog/actions-architecture/markdown) - [Why Most Healthcare AI Can't Handle Real Patient Conversations](https://www.amigo.ai/blog/dynamic-behaviors): Amigo's Dynamic Behaviors allow AI agents to recognize what matters during clinical conversations and adapt in real time. Published 2025-12-17; author: Claire Uhm. [Markdown](https://www.amigo.ai/blog/dynamic-behaviors/markdown) - [Amigo Establishes Medical Advisory Board and Appoints Dr. Jay Shah as Chief Medical Advisor](https://www.amigo.ai/blog/amigo-establishes-medical-advisory-board): Stanford Health Care's Chief of Medical Staff to Guide Clinical Strategy for Healthcare AI Platform Published 2025-10-30; author: Richard Wang. [Markdown](https://www.amigo.ai/blog/amigo-establishes-medical-advisory-board/markdown) - [The path to patient-facing AI? Follow Waymo's lead](https://www.amigo.ai/blog/the-path-to-patient-facing-ai): 5 principles for patient-facing AI, according to top clinicians Published 2025-10-29; author: Christina Farr & Anjalee Khemlani. [Markdown](https://www.amigo.ai/blog/the-path-to-patient-facing-ai/markdown) - [Deep Dive: Context Graphs](https://www.amigo.ai/blog/context-graphs): How Amigo's Context Graph architecture enables structured yet flexible clinical conversations. Published 2025-10-08; author: Ali Khokhar. [Markdown](https://www.amigo.ai/blog/context-graphs/markdown) - [Amigo Deep Dive: Digital Health Wire](https://www.amigo.ai/blog/digital-health-wire): AI moves fast, but trust moves slow. That's why Digital Health Wire is launching a new series to spotlight the companies taking AI from promise to practice. Published 2025-09-23; author: Jason Barry. [Markdown](https://www.amigo.ai/blog/digital-health-wire/markdown) - [AI Doctors Are Like Self-Driving Cars: Lessons from Waymo to Build Trust](https://www.amigo.ai/blog/lessons-from-waymo): Building trust through scoped domains, rigorous simulation, and human-in-the-loop safeguards. Published 2025-09-19; author: Richard Wang. [Markdown](https://www.amigo.ai/blog/lessons-from-waymo/markdown) - [Deep Dive: Functional Memory](https://www.amigo.ai/blog/functional-memory): Introducing Amigo's Functional Memory architecture that allows agents to think, learn, and remember like healthcare professionals. Published 2025-09-10; author: Ali Khokhar. [Markdown](https://www.amigo.ai/blog/functional-memory/markdown) - [Healthcare AI Infrastructure: Build vs. Buy?](https://www.amigo.ai/blog/build-vs-buy): Navigating implementation trade-offs, hidden costs, and architectural complexity in clinical AI deployment. Published 2025-08-21; author: Ali Khokhar. [Markdown](https://www.amigo.ai/blog/build-vs-buy/markdown) - [Beyond Benchmarks: Why Healthcare AI Needs Real-World Validation](https://www.amigo.ai/blog/beyond-benchmarks): Standardized performance benchmarking fails to prepare AI agents for the nuanced situations faced by actual patient populations. Published 2025-07-30; author: Ali Khokhar. [Markdown](https://www.amigo.ai/blog/beyond-benchmarks/markdown) - [Amigo Partners with Eucalyptus to Scale AI-Powered Healthcare Delivery Across Global Telehealth Network](https://www.amigo.ai/blog/amigo-partners-with-eucalyptus): Strategic partnership aims to deploy AI health assistants across Australia's largest telehealth platform serving 200,000+ patients in four countries. Published 2025-07-28; author: Richard Wang. [Markdown](https://www.amigo.ai/blog/amigo-partners-with-eucalyptus/markdown) - [Amigo Deep Dive: Healthcare AI Guy](https://www.amigo.ai/blog/healthcare-ai-guy): Healthcare AI Guy interviewed Ali Khokhar, Amigo's CEO, about how they're using simulation to build trust, turning agents into action-takers, and redefining what build vs. buy means in healthcare AI. Published 2025-07-17; author: Healthcare AI Guy. [Markdown](https://www.amigo.ai/blog/healthcare-ai-guy/markdown) - [Preventing Clinician Burnout through AI-Powered Care](https://www.amigo.ai/blog/preventing-clinician-burnout-through-ai-powered-care): How intelligent patient engagement transforms the care experience from first contact to clinical outcomes. Published 2025-06-03; author: Ali Khokhar. [Markdown](https://www.amigo.ai/blog/preventing-clinician-burnout-through-ai-powered-care/markdown) - [Evaluations as the Path to Trust](https://www.amigo.ai/blog/evaluations-as-the-path-to-trust): How we designed an evaluations system to achieve 99.9% safety scores in high-stakes healthcare environments. Published 2025-05-28; author: Ali Khokhar. [Markdown](https://www.amigo.ai/blog/evaluations-as-the-path-to-trust/markdown) - [How Amigo Solves AI's Trust Crisis](https://www.amigo.ai/blog/how-amigo-solves-ais-trust-crisis): The trust problem prevents AI adoption in critical systems. Amigo builds responsible AI that performs reliably in high-stakes environments. Published 2025-05-06; author: Ali Khokhar. [Markdown](https://www.amigo.ai/blog/how-amigo-solves-ais-trust-crisis/markdown) ## Machine-readable resources - [Full published article text](https://www.amigo.ai/llms-full.txt): Markdown with article URLs, authors, and dates - [Blog index](https://www.amigo.ai/data/blog-index.json): Article metadata and Markdown URLs - [Sitemap](https://www.amigo.ai/sitemap.xml) - [RSS feed](https://www.amigo.ai/feed.xml) - [Atom feed](https://www.amigo.ai/feed.atom) - [JSON feed](https://www.amigo.ai/feed.json) ## Contact and documentation - [Book a demo](https://www.amigo.ai/book-demo) - [Platform documentation](https://docs.amigo.ai/getting-started/amigo-overview) - [Trust center](https://trust.amigo.ai/) - [System status](https://status.amigo.ai/) - Email: contact@amigo.ai ## Optional - [Careers](https://www.amigo.ai/careers) - [Student Researcher](https://www.amigo.ai/student-researcher) - [Privacy](https://www.amigo.ai/privacy) - [Terms Of Service](https://www.amigo.ai/terms-of-service) --- ## Full published articles The following text comes from the published blog. Each article includes its source URL and publication date. Customer results and product announcements describe their published context. --- # Why I'm Joining Amigo as Chief Medical Officer Source: https://www.amigo.ai/blog/dr-jay-shah-amigo-chief-medical-officer Author: Jay Shah, MD Published: 2026-06-08 I became a cancer surgeon to help people through their hardest days. I'm joining Amigo so I can reach more people and multiply my impact. Today, I’m stepping into a new role as Amigo's Chief Medical Officer where I’ll lead the clinical strategy that guides how our AI is built and deployed. I’ve spent the last quarter century in operating rooms and at the bedside. In my work as a cancer surgeon and Chief of Medical Staff at Stanford Health Care, I have learned just how central trust is to the entire healthcare enterprise. While we get into medical school because of our intelligence, it is our ability to form trusting relationships that truly defines our success as physicians. These relationships - whether between the patient and their physician, the care team and their organization, or medicine and society - all ultimately rest on a bedrock of trust. Now, as we usher in the age of AI, I firmly carry forward the conviction that any technology that claims to improve healthcare delivery also must earn the trust of patients and clinicians. With new healthcare AI startups launching at a fever pitch, I hold significant concern that the long-treasured element of trust in medicine may be swept aside in the name of purported efficiency gains and valuation end runs. When I started advising Amigo in 2025, I came prepared to stress-test every claim. I’ve seen healthcare technology companies promise more than they could deliver, with solutions that physicians wouldn’t feel comfortable standing behind. I asked Ali and John hard questions about how Amigo agents are validated and what happens when cases fall outside of what the model expects. I was pleased to discover that clinical rigor and safety are the lynchpins of everything Amigo does. The fact that the entire Amigo family carries a deep humility about the weight of building in healthcare speaks directly to my heart. The next decade of medicine will be defined by how responsibly we bring intelligence into care. Done irresponsibly, AI will erode the trust that is the cornerstone of everything we do in medicine. If we do it right, AI can become the final keystone in the house of medicine that will allow billions more people the access to timely, high-quality medical care required for human flourishing. My core motivation for serving as the Chief Medical Officer at Amigo is to help lead society down this more responsible path. In my role as CMO, I’ll be working with our engineers and customers to ensure every clinical agent we deploy is grounded in evidence and shaped by the clinical minds who understand the care it supports. I also plan to keep one foot in the clinical world by continuing to practice as a cancer surgeon to ensure the questions we answer stay tied to what matters at the bedside. I’m not leaving medicine to become a tech bro; I’m looking around the corner to see where medicine is heading and trying to clear the path of hurdles and false idols. Meeting patients at what is often the darkest moment of their lives and having them trust me to walk with them on their cancer journey has been among the greatest honors of my life. And now, helping figure out how to honor and preserve that hard-earned trust as we learn to responsibly weave AI into healthcare seems like the natural continuation of this work. I feel grateful, excited, and a little bit scared to undertake this task. I aim to lead medicine into the arena where it can help the most people thrive. If you have ideas that can help me succeed in this endeavor or if you’re thinking about how best to leverage AI to help your patients and clinicians, I welcome your thoughts and your feelings. ### About Dr. Jay Shah *Dr. Jay Shah is Chief Medical Officer at Amigo. A cancer surgeon and associate professor of Urology at the Stanford University School of Medicine, he was the Chief of Medical Staff at Stanford Health Care and is a nationally recognized expert in robotic surgery and bladder cancer treatment.* *Prior, he served as Center Medical Director for the Genitourinary Center at MD Anderson Cancer Center, where he launched the bladder cancer robotics program and developed an enhanced recovery program for patients undergoing bladder removal surgery.* *A graduate of Harvard College, he completed his medical degree and residency at Columbia University, where he was elected to the Alpha Omega Alpha Medical Honor Society. He has been named Physician of the Year and received the Gold Foundation Excellence in Teaching Award for his work mentoring the next generation of physicians.* --- # Too Many Apps, Not Enough Medicine: The Platform Approach Source: https://www.amigo.ai/blog/amigo-platform-approach Author: Ross Green, MD Published: 2026-05-08 Why healthcare's digital fragmentation is a patient safety crisis, and what a platform can do about it ## Abstract (TL;DR) *Healthcare has a fragmentation problem: clinicians log into an average of 12 different systems just to do their jobs, none of which communicate with each other. This is far more harmful than simply being inefficient since it drives burnout, errors (with many asking the logical question: "Am I getting the full clinical picture or just bits and pieces?"), and billions in wasted spending, let alone how it drives practitioners and staff alike mad.* *The solution is definitely not more apps, but rather a platform approach. Amigo is positioned as a unified AI operating system that replaces the fragmented stack with a single environment where purpose-built swarms of agents collaborate and handle everything from documentation to prior authorization to chronic disease monitoring, and beyond. Critically, each agent goes through a physician-guided validation process, or an "AI residency," before it ever touches a clinical setting. This addresses the core reason that past health tech (the EHR being the prime example) has “failed” according to many prominent observers, since they involve deployment without adequate clinical validation.* *The bottom line: medicine deserves what every other high-stakes field eventually builds, which is an all-in-one coherent system, designed around the people using it, that lets clinicians and healthcare operators focus on the patient in front of them rather than the login screen.* ## **Introduction** In medicine, we are taught early that a scattered history is a dangerous history. In this light, when a patient can't tell you their medications, their allergies, or the name of their last specialist, we slow down, we triangulate, and we (understandably) worry that we’re missing something. In short, it is the incomplete picture where errors often live. It is ironic, then, that we have spent the last two decades building a digital infrastructure for healthcare that is, by design, scattered. To be sure, there are some areas of this fragmentation that clinicians simply have no control over, such as issues of various EHR systems at different healthcare organizations being unable to “talk” to one another (even if using the same company, such as Epic or Oracle). However, what healthcare companies and providers/clinicians can control is what tools they use, and the result has been that these organizations have generally chosen, on purpose or not, a fragmented digital ecosystem where every problem has an app, every app has its own login, and none of them talk to each other. This is precisely where the idea of a platform, or one cohesive operating system, arises. ## **The Proliferation Problem** Ask any clinician how many systems they log into on a given day. A recent industry analysis found that clinicians navigate an average of 12 different systems and applications just to access current patient records [1]. There is the EHR for documentation, a separate portal for imaging, another for labs, a scheduling tool, a secure messaging platform, a patient engagement app, a prior authorization portal, a care management dashboard, a telehealth interface, and on it goes. Perhaps most critically, as a 2025 systematic review in *Information* documented, current interoperability standards like FHIR cannot reliably retrieve patient records stored across multiple systems with diverse implementation guides, since the standards were designed for institutional exchange rather than a unified clinician experience [2]. The consequences extend well beyond inconvenience. Poor interoperability is estimated to cost the U.S. healthcare system over $30 billion annually in avoidable inefficiencies, including administrative overhead, unnecessary testing, and delayed treatment decisions [3]. As authors published in the *New England Journal of Medicine* noted as recently as February 2026, despite years of federal effort, including TEFCA, FHIR mandates, and CMS' Digital Health Ecosystem initiative, fundamental economic and policy barriers to true interoperability remain stubbornly in place [4]. A fragmented data environment isn't just inefficient. It is, in the truest clinical sense, a liability. A 2025 study in the *European Journal of Public Health*, drawing on data from 9,526 primary care physicians across 10 OECD countries, found that digital health tools, which are intended to ease burdens, paradoxically contribute to burnout when they fragment rather than streamline workflows [5]. Meanwhile, the American Medical Association (AMA) has documented that EHR-related burdens, including excessive inbox volume, workflow interruptions, and poor interoperability, are among the most consistent drivers of physician burnout [6]. As of 2026, more than half of U.S. clinicians report symptoms of burnout, and the digital environment they work in is a recognized contributor [7]. ## **A Platform, Not a Portfolio of Apps** We need to take control of the things we can control. Perhaps we cannot change how EHR interoperability works, or how these EHR systems “talk to each other,” as this would require governmental regulation, and who is realistically going to wait for that to happen, if it happens at all? As such, the answer to fragmentation is consolidation. Not just simply of data ownership, but of intelligent, interoperable workflows and tools. This is the premise behind Amigo, a multi-agent AI operating system platform designed to span the breadth of healthcare's most demanding use cases within a single, unified platform. Rather than deploying a separate tool for scheduling, another for prior authorization, another for clinical documentation, another for patient follow-up, and yet another for population health monitoring, Amigo coordinates purpose-built AI agents across all of these domains simultaneously. The clinical picture that emerges is, by design, complete. As such, a physician using Amigo isn't switching contexts or toggling between windows; rather, they are operating within a coherent environment where the agents work in concert, surfacing the right information at the right moment for the right decision, and thereby ensuring as complete a clinical picture as possible is reflected in the agent’s decisions. The breadth of what Amigo covers is far from incidental; rather, it is the very point. Across clinical documentation, care coordination, revenue cycle management, patient engagement, real-time clinical decision support, prior authorization, appointment scheduling, referral management, and chronic disease monitoring, the platform is designed to replace a fragmented stack of siloed tools with a single operating layer. The physician who once needed five logins before noon now needs one. And on the issue of lack of interoperability between different EHRs: with Amigo, the platform can “connect” (using APIs) these various data sources under one roof, thereby allowing as much of your patient’s data to be utilized by Amigo’s operating system tools. ## **Trust Is the Feature, And It Is Not The Afterthought** The most consequential question in clinical AI is trust and capability. A system that can generate a note or flag a drug interaction is only as valuable as the confidence a clinician has in its outputs. And this is not a small concern. In this vein, the healthcare landscape is littered with AI tools that were technically impressive, poorly validated, and quietly abandoned after generating alarm fatigue, erroneous recommendations, or a lack of trust around the output because it may not include the entire clinical picture. Amigo addresses this through what might be called an AI residency model, or a deliberate, structured training and validation process for each agent before it ever operates in a clinical setting. Just as medical residency fosters collaboration across specialties, Amigo's platform is built for cross-functional coordination, with every AI tool operating within a unified system rather than in isolation. Central to this approach is a rigorous validation process in which active clinicians assess each agent's performance against real-world clinical scenarios prior to any deployment. The process goes further by leveraging digital cloning to simulate both common practice challenges and low-frequency edge cases at scale, exposing agents to millions of scenarios they may one day encounter in the field. Intelligent AI judges play the role of attending physicians in these simulated scenarios, providing feedback on what the agent could do better next time. This process is completed for *all* agents built on Amigo’s platform, ensuring the same high bar for safety and accuracy across every single patient workflow. The stakes of getting this wrong are far from hypothetical. A 2026 commentary in *NEJM Catalyst,* reflecting on two decades of digital health disappointment, argued that the electronic health record, despite near-universal adoption, “failed” to deliver on its promise not because the technology was absent but because tools were deployed without sufficient organizational and clinical validation, generating administrative burden rather than clinical value [8]. The lesson is not subtle: technology that physicians don't trust doesn't get used, and technology that gets used without physician validation doesn't deserve to be. ## **The Case for Consolidation** The argument for a unified platform over a portfolio of disconnected apps is ultimately the same argument we make in clinical medicine every day: context matters, and context requires continuity. In this light, a physician making a prescribing decision needs to know the patient's kidney function, their current medications, their insurance coverage, and the last time a relevant lab was drawn, all simultaneously, not across four separate logins with tools/systems that simply don’t communicate. Similarly, a care coordinator managing a post-discharge patient needs the discharge summary, the follow-up appointment, the pharmacy status, and the patient's response to an outreach message, all in one view rather than five disparate systems. Amigo is not another app. It is the argument that medicine deserves a platform, one that was built with physicians, validated by physicians, and designed to restore what fragmentation has quietly taken from the practice of medicine: the clarity to focus on the patient in front of you. ### **References** [1] Hart Health. Solving Fragmented Healthcare Data with Interoperability. Hart.com. September 10, 2025. Available at:[ https://hart.com/blog/how-interoperability-can-solve-fragmented-healthcare-data-challenges](https://hart.com/blog/how-interoperability-can-solve-fragmented-healthcare-data-challenges) (accessed May 7, 2026). [2] Jendly M, et al. From Data Silos to Health Records Without Borders: A Systematic Survey on Patient-Centered Data Interoperability.* Information *(MDPI). 2025;16(2):106. doi:10.3390/info16020106 [3] West Health Institute. *The Value of Medical Device Interoperability: Improving Patient Care with More Than $30 Billion in Annual Health Care Savings.* March 2013. Available at: [westhealth.org/wp-content/uploads/2015/02/The-Value-of-Medical-Device-Interoperability.pdf](http://westhealth.org/wp-content/uploads/2015/02/The-Value-of-Medical-Device-Interoperability.pdf) (accessed May 7, 2026) [4] Halamka JD, Tripathi M. The Next Chapter in Health Care Interoperability. *N Engl J Med*. February 7, 2026. doi:10.1056/NEJM p2511798 [5] Jendly M, Santschi V, Tancredi S, et al. Primary care physician digital health profile and burnout: an international cross-sectional study. *Eur J Public Health. *2025 Jul 15:ckaf106. doi:10.1093/eurpub/ckaf106 [6] American Medical Association. Electronic Health Record (EHR) Use Research. *American Medical Association*. Updated March 2026. Available at: [https://www.ama-assn.org/practice-management/digital-health/electronic-health-record-ehr-use-research](https://www.ama-assn.org/practice-management/digital-health/electronic-health-record-ehr-use-research) (accessed May 7, 2026). [7] Virginia Center for Health Innovation. APP Burnout in Primary Care 2026. February 9, 2026. Available at:[ https://www.vahealthinnovation.org/virginia-joy-in-healthcare/app-burnout-in-primary-care-2026/](https://www.vahealthinnovation.org/virginia-joy-in-healthcare/app-burnout-in-primary-care-2026/) (accessed May 7, 2026) [8] Wachter R, Lee TH. Beyond the Hype: How AI Is Finally Delivering on Digital Health's Promise. *NEJM Catalyst*. February 8, 2026. doi:10.1056/CAT.26.0043 --- # How We Build Self-Improving Agents Source: https://www.amigo.ai/blog/agent-forge Author: Claire Uhm Published: 2026-04-16 Our answer to the post-launch optimization problem The barrier to building AI agents has reduced dramatically over the past year. You can now describe what you want to build in plain language and have a working agent prototype within the same day. The shared thesis across the industry is that agent creation should be effortless, and we agree. Lowering the barrier to building agents means organizations can discover new ways to improve their workflows without committing engineering resources upfront. The natural result is that more agents are making it to production, but the industry hasn't made the same progress on what happens after launch. Agents need to be continuously monitored and improved to ensure their performance doesn’t degrade over time, and the ongoing work of optimizing agents is just as engineering-intensive as building them used to be. Our solution to this is the *Agent Forge*, our agent development platform where AI agents handle that work end-to-end nearly autonomously. ## Self-Improving Agents Amigo's clinical agents are self-improving by design. The *Agent Forge* makes this possible by powering a separate team of AI coding agents that automatically detect issues, build fixes, and prepare updates for human review. ![How Agent Engineers and coding agents in the *Agent Forge* collaborate to improve Amigo’s clinical agents](https://www.amigo.ai/images/blog/content/agent-forge-0.png) The process begins with one of our agent engineers describing the changes they want to make, which prompts coding agents to pull the clinical agent configuration, make the edits, validate that they work by running simulations against thousands of AI-simulated patient personas, check for regressions, and prepare the updates for human review. Once that process is complete, our engineers approve or reject the changes, ensuring humans always have the final say. The loop closes with the coding agents implementing any approved changes to the Amigo clinical agent. This cycle allows our agents to continuously learn and improve, with built-in checks and balances to prevent unwanted changes from reaching patients. ## Testing Beyond Pre-Defined Scenarios Before an Amigo agent talks to a real patient, it goes through millions of simulated conversations designed to test the edges of its capabilities. Several other agent platforms on the market offer some form of testing, but the *Agent Forge* treats agent development as a verification problem. The difference is that **testing** checks whether an agent can handle a set of pre-defined scenarios, whereas **verification** systematically finds scenarios outside of the problem set. ![Standard testing covers pre-defined scenarios, whereas verification systematically discovers new conversation paths](https://www.amigo.ai/images/blog/content/agent-forge-1.png) The *Agent Forge*'s simulation engine automatically discovers conversation paths the agent has never encountered to deepen coverage in areas where quality is weakest. It deliberately seeks out failure modes by running every scenario against simulated patients with different temperaments, because an agent that performs well with a calm and articulate patient may respond differently when that patient is anxious, skeptical, confused, or frustrated. The goal is to prove the agent can handle the full range of real-world interactions. ## **Measurement That Resists Gaming** Most agent platforms optimize toward a single metric, such as CSAT (customer satisfaction score), task completion, or resolution rate. The problem with this is well understood. When a measure is treated as a target, it stops being a good measure. For example, an agent optimized for CSAT learns to become agreeable rather than accurate, and an agent optimized for resolution rate learns to rush through conversations before gathering sufficient context. Metrics may appear to improve on the surface, but the patient experience suffers. Rather than optimizing for any single number, the *Agent Forge* measures agent performance across multiple dimensions simultaneously, and an interaction has to perform well across all of them to count as a success. Improvement in accuracy at the cost of empathy is a failure. So is an increase in resolution speed if it leaves patients feeling confused. ![The *Agent Forge* measures agent performance across multiple dimensions simultaneously](https://www.amigo.ai/images/blog/content/agent-forge-2.png) ## **Catching Drift Before It Compounds** Agents degrade over time as underlying language models get updated or shifting patient populations and seasonal patterns change the mix of incoming questions. An agent that was verified at launch may begin behaving differently, often in subtle ways that go unnoticed until a pattern of failures has already built up. The *Agent Forge* continuously monitors for this kind of drift and is able to distinguish between two different types. **Performance drift** is when the agent's outputs begin to change, whereas **dimensional drift** is when the problem space itself has changed and the original agent is no longer able to cover it effectively. Each requires a different response, and the *Agent Forge* is able to initiate a structured cycle to understand what changed, determine the right fix, validate it through simulation, and deploy it with a human in the loop. ## **The Partnership Model** Many agent platforms are designed to abstract customers away from complexity. “Upload your SOPs, describe what you want, and the platform handles the rest.” This black box approach is good enough for most industries, but in healthcare, it’s important for the clinical team to stay involved. The people closest to the clinical workflows are the ones who understand what a good interaction looks like for their specific population and protocols. The *Agent Forge* is built for that kind of shared ownership, where the healthcare organization contributes their expertise on what good care looks like, and Amigo turns that input into an agent that delivers it reliably. The *Agent Forge* is the interface between those two responsibilities, and as customer teams familiarize themselves with the platform, they can even begin driving agent development themselves. ## **Get Started with Amigo** If you're a healthcare leader evaluating clinical agents, it’s important to understand what happens after launch. Your agents will need to keep up as edge cases accumulate and conditions change. It’s why we built the *Agent Forge*, and we'd love to show you how it works. [Book a demo](https://www.amigo.ai/book-demo) to see Amigo agents in action. --- # Amigo AI Raises $11M Series A to Train Clinical AI Agents Like Doctors Source: https://www.amigo.ai/blog/amigo-series-a Author: Ali Khokhar Published: 2026-03-10 This brings Amigo's total funding raised to $17M. [Watch the video](https://www.youtube.com/watch?v=HiM5xaJp_QE) My mom was diagnosed with breast cancer when I was eight years old. Over the next six years, through two diagnoses, I watched her navigate a healthcare system that was never built to support her. What I remember most is how much work it was just to be sick. The same story told to every new specialist, the same exams repeated, referrals made and forgotten. And through all of the chemo, side-effects, and hard nights, the burden of coordinating my mom’s care always fell on our family. She passed away when I was fourteen. What I couldn't shake was the feeling that so much of it didn't have to be that hard. That was over fifteen years ago. Medical research has come a long way, but the science has advanced faster than the system built to deliver it. Now, the bottleneck is human capacity, and the structural problems that prevent patients from getting care are worse than ever. 100 million Americans don't have access to a primary care provider. The average wait time to see a physician is 31 days. By 2030, the world will face a shortage of 11 million health workers. I grew up around clinicians and am surrounded by them still, and I've watched them face this impossible reality firsthand: burnout, waitlists that stretch for months, and patient rosters that outpace even the most dedicated physicians. AI in healthcare fundamentally changes the equation. But for AI to solve the clinical capacity crisis, we knew two things had to be true: **1. AI that actually delivers care.** Most healthcare AI today focuses on relieving administrative burden and helping clinicians work faster, but it can’t fill in when there aren't enough clinicians to go around. That's a fundamentally different problem that no amount of scheduling and note-taking automation can solve, and it requires intelligent AI agents that can actually handle high-value clinical workflows like intake and triage, chronic care management, and 24/7 virtual care. Our clinical agents help healthcare organizations deliver care at a scale their human teams could never reach alone. **2. AI that is safe enough to earn that responsibility.** If AI is going to operate in this space, it has to be held to the same standard as the clinicians it works alongside. To achieve this, we train our agents like doctors. Before any Amigo agent interacts with a real patient, it goes through what we call the **Digital Residency:** millions of simulated patient scenarios, modeled on the actual population of each healthcare organization it will serve. Those simulations are deliberately adversarial: edge cases, ambiguous presentations, distressed patients. Agents continue to learn and improve across dimensions like clinical accuracy, empathy, and harm prevention until they reach a 100% safety pass rate. Every agent shares a unified patient context in real time, escalates automatically, and hands off cleanly across the full arc of a patient's care journey. In the last six months alone, Amigo agents have completed over 3 million patient encounters around the world with zero safety incidents. The Digital Residency is what makes that possible. It's how we earn the trust of every organization we work with, and it's the foundation everything else at Amigo is built on. Today, I’m incredibly proud that Amigo powers clinical AI for leading healthcare organizations around the world, including Eucalyptus, Diverge Health, and The Care Clinic. Our agents operate in over 100 languages with native integrations into all major EHRs. To support our next chapter, we've raised $17M in total funding, including an $11M Series A led by Madrona with participation from Optum Ventures, and a seed round co-led by General Catalyst and GSV Ventures. Our mission is to make high-quality healthcare available to everyone in the world, and we believe safe and reliable clinical agents are the only way to make this possible. Our commitment to every organization we partner with is to be the single platform your care team grows on, where every agent you’ll ever need is trained, deployed, and working together in one place. The system my mom navigated was broken. We have a real chance to build something better together. If you're a healthcare organization thinking about how AI can help your care team reach more patients safely, I'd love to talk. *Reach me directly at ali@amigo.ai.* --- # How Clinical Agents Get Built and Trained Source: https://www.amigo.ai/blog/how-clinical-agents-get-built-and-trained Author: Nikhil Krishnan Published: 2026-02-27 A walk through of what agents are and how Amigo builds them *The following is a direct transcription of the article. To read this post on Out-Of-Pocket Health's website, please visit* [***this link***](https://www.outofpocket.health/p/how-clinical-agents-get-built-and-trained?utm_source=sponsored&utm_medium=partner&utm_campaign=q1_oop&utm_content=amigo_blog)*.* ## **TL:DR** Everyone is talking about AI agents in healthcare. Today, we’ll talk about what agents are and how they compare to types of automation in the past. There are lots of different ways to train, test, and deploy these agents. We’ll also cover [Amigo](https://www.amigo.ai/?utm_source=sponsored&utm_medium=newsletter&utm_campaign=q1_oop), a company that enables providers to build their own agents. We’ll go through the product itself and how they create provider-specific simulated environments to customize and battle test agents. The company is taking the approach of building a platform to create multiple agents together vs. selling pre-made agents. We’ll talk about the pros and cons of their approach, and the tradeoffs they’re making. ‍*This is a sponsored post. You can read more about my rules/thoughts on sponsored posts* [*here*](https://outofpocket.health/p/an-update-about-out-of-pocket)*.* ## **Company Name - Amigo** ![Blog image](https://www.amigo.ai/images/blog/content/how-clinical-agents-get-built-and-trained-0.png) Amigo has built a platform to build AI agents specific to your practice/clinic. We’ll talk about how they do that in a second. I got a 2 on AP Spanish but even I know “Amigo” means “friend”. This would be like me starting an AI company and calling it “The Boyz”. ## **What is an Agent?** We’ve had automation in healthcare for quite a long time. It might be helpful to think about an agent compared to other types of automation you might have heard about in the past. Let’s start with **Robotic Process Automation**. Companies map out workflows a human does - things like clicking on screens, copying and pasting data, etc- and a computer repeats that workflow exactly. This automation exercises no judgment, has a very specific “track” to follow, and frequently breaks if small changes happen in that flow (e.g. a button changes place). It’s brittle and also very narrow in what it can do. Then we had **chatbots**. If AI did that “post something from 2016 trend” they’d post a chatbot. Chatbots are a bit more evolved in that they have branching decision trees in the background, so they could handle a wider variety of inputs. However, as soon as they deviated from the path, they broke and required escalation. Plus, there was no chat history so each conversation was brand new, like Memento for chat. There’s nothing more fun than going to your favorite telecom carrier, being told to use the chatbot with preloaded options, and then having it fail 90% of the time so you need to talk to someone anyway. Today, we come to **agents**. Agents are a type of automation that uses this newfangled AI to power it, but they have a few components that make them extra powerful. ![Blog image](https://www.amigo.ai/images/blog/content/how-clinical-agents-get-built-and-trained-1.png) 1. **Differing Personalities** - You have the ability to give different agents different personalities and role types. This allows you to embed different rules, aggressiveness, tone, etc. into a given automation. This matters because it will shape how the underlying language model interprets info, responds, and escalates. 2. **Context Graphs** - Context graphs give information on HOW decisions are made and the context needed to make those decisions. For example, when a patient asks about their rash the agent needs context about the rest of the patient's condition, what questions to ask in that situation, and what escalation should be, based on some general guidance (with flexibility to adapt). When a doctor sees that message, a similar mental framework is kicking in, developed from seeing that scenario many times and knowing the rest of the patient’s health history. 3. **Actions** — The agent is allowed to interact and push/pull data to different systems based on the permissions it has. The breadth of what they can do tends to be larger as well, because they can go through API endpoints or use browsers just like how you would use a computer. Today, you even see companies that have specific pathways for agents to interact with their system (e.g. [Model Context Protocol](https://modelcontextprotocol.io/docs/getting-started/intro)). 4. **Memory** - Agents have memory. This is not only remembering previous times you’ve interacted with them but also how that agent has interacted with different systems and users that might be relevant. Because of the vast amounts of data that could count as memory, an entire hierarchy has to be built that figures out which memory is important for a given task or context. Companies tend to build with very strong opinions about what memory is relevant to pull at which times. ![Blog image](https://www.amigo.ai/images/blog/content/how-clinical-agents-get-built-and-trained-2.png) The result is that agents have opinions on how to get things done and flexibility to attempt other things if things don’t go correctly. They have dynamic behavior; if something changes from the given path, they can change with it. When done correctly, they basically function like employees at your job. ## **What Pain Point is Being Solved? What Does Amigo Do?** What’s hard about deploying agents into clinics is that every clinic is extremely different. They have different rules, see different kinds of patients, have different points of view on triaging, etc. So to learn context and everything that makes agents powerful is really hard. This is especially true when it comes to patient facing agents or agents that need to handle any level of clinical task. The prior authorization flow will look similar between a rural surgery center and a metropolitan gastroenterologist, but the clinical and patient facing tasks will look very different. The risks are higher and need to be more specific to the setting it’s deployed in, so the agents need to be higher fidelity. ![Blog image](https://www.amigo.ai/images/blog/content/how-clinical-agents-get-built-and-trained-3.png) Amigo has a few components to do this. The first is onsite deployment + an agent factory. They send people to ingest a ton of data from a practice EHR, practice management system, scribes, standard operating procedures, and interviews with the practice managers. By doing this they: 1. Build profiles of the different types of patients a practice sees. 2. Learn the rules of interacting with patients. Both the explicit rules that are written down, and the implicit rules that come from interactions. 3. Learn the personality types to give different agents, what actions to do in different scenarios, and what memory to pull, when. As you can imagine, this is a massive amount of data in itself, so Amigo has a ton of their own agents that do the work of ingesting a lot of this raw data and creating that foundation. The next step is to take the agents that were built and put them in a simulated world with millions of potential patient encounters. It’s like when you’re thinking about a fake argument in your head in the shower and the cool things you’ll say to win, but for AI. Those agents are then judged by another AI that objectively tells them what went wrong, self-analyzes, and makes improvements. This loop helps the agent improve and figure out cases where it won’t be as strong (and therefore may need to escalate to a human in the loop). With this, they now have agents that are actually usable in patient facing encounters. You can build different kinds of agents for different tasks on top of this foundation, and then they have the tools to monitor how the agent is performing. The Amigo CEO called this the “AI residency program”. AI can pass the MD licensing exams with flying colors, but the doctors learn the real-world implementation when they become residents at a hospital or clinic. The only way that happens is by seeing things a million times and learning from mistakes, which an agent does in your practice-specific simulated environment. And like real residents, you don’t need to pay them much…too real? ![Blog image](https://www.amigo.ai/images/blog/content/how-clinical-agents-get-built-and-trained-4.png) ## **An Agent Example - Side Effect Management** A common use case that gets built with these clinical agents is side effect management. Let's say you're a digital health company prescribing GLP-1s for weight loss. You’re very original. Patients are texting constantly about side effects. They’re nauseous, the injection site has weird reactions, they didn’t use the pen properly, etc.. They want answers at odd hours and it’s not really worth the time for the practice to have someone clinical answer questions about stool consistency or “you up? wyd”. Amigo works with the clinical team to build the agent's brain. They define: - The personality - Warm but direct, avoids medical jargon. - The context graph - States for symptom assessment, severity determination, checking medication history, providing guidance, and escalating. - The behaviors that override - For example, "if a patient mentions chest pain or severe dehydration, escalate immediately, regardless of where you are in the conversation.” - What memory dimensions matter here - medications, allergies, how long they've been on the GLP-1, past side effects they’ve had, past dosages, etc. The clinical team sets the success metrics: empathy needs to be above an 8, clinical accuracy higher, and appropriate escalation has to be 100% of cases that need it. Amigo creates thousands of synthetic patients that mirror the practice's actual population and runs the agent through scenarios: "Patient reports mild nausea after first injection"... "Patient has been vomiting for 3 days and can't keep water down"... "Patient is talking about proteinmaxxing but I have no idea what they’re talking about." Then the AI Judge evaluates each conversation. The transcript, internal reasoning for the agent, and the outcome. Did it check how long the patient has been on the medication? Did the right dynamic behavior fire when the patient mentioned they couldn't keep fluids down? An Amigo engineer reviews, approves, runs it again. This loop continues until it hits thresholds across all metrics. ![Blog image](https://www.amigo.ai/images/blog/content/how-clinical-agents-get-built-and-trained-5.png) The agent then gets deployed. Patients can ask questions about side effects they’re feeling and get answers, a new appointment, or get switched to a new drug (with a doc reviewing, async). The clinical team gets dashboards showing where the agent is strong and where it's struggling. The founders can tell their investors they’re AI-enabled care delivery. New edge cases from real patients feed back into the simulation environment. This process is done for every agent type, clinic, etc. ![Blog image](https://www.amigo.ai/images/blog/content/how-clinical-agents-get-built-and-trained-6.png) ## **What Is The Business Model And Who Is The End User?** Amigo charges a base platform fee plus usage-based pricing. Customers get: - Access to the platform (agent creation, simulation/testing environment, monitoring) - A forward-deployed "Agent Engineer" embedded with their team. Oh god there are subspecialties of forward deployed engineers now. - Compute - A friend They have two groups of customers. The first are digital health companies doing care delivery. Customer engagement increases and the patients that engage with the patients tend to have higher lifetime values. The second is traditional providers. Traditional clinics are using these agents to handle patient-facing clinical tasks that might be bottlenecked by clinical hiring (e.g. nurse phone lines, triaging, chasing down care gaps, etc.). This helps practices scale without needing to hire more. They work across a range of specialties including primary care, orthopedics, behavioral health, women's health, and wellness/longevity. ## **Job Openings** Amigo is looking to triple their team in 2026. They're hiring for: - [Agent Deployment Strategist](https://www.amigo.ai/careers/agent-deployment-strategist-b4caea14-54af-4b8a-a101-efe593cad674?ashby_jid=b4caea14-54af-4b8a-a101-efe593cad674&utm_source=sponsored&utm_medium=newsletter&utm_campaign=q1_oop) - [Staff Software Engineer (Infra)](https://www.amigo.ai/careers/staff-software-engineer-data-427325d1-3b63-47e5-99df-1d7ee482bb0c?ashby_jid=427325d1-3b63-47e5-99df-1d7ee482bb0c&utm_source=sponsored&utm_medium=newsletter&utm_campaign=q1_oop) - [Staff Software Engineer (Backend)](https://www.amigo.ai/careers/staff-software-engineer-backend-32c5d071-c041-48c0-9592-7161c6ce23b2?ashby_jid=32c5d071-c041-48c0-9592-7161c6ce23b2&utm_source=sponsored&utm_medium=newsletter&utm_campaign=q1_oop) - [Staff Full-Stack Engineer](https://www.amigo.ai/careers/staff-full-stack-engineer-d4a79eec-1304-4a94-981a-146b5e4da9c3?ashby_jid=d4a79eec-1304-4a94-981a-146b5e4da9c3&utm_source=sponsored&utm_medium=newsletter&utm_campaign=q1_oop) - [Account Executive](https://www.amigo.ai/careers/founding-enterprise-account-executive-traditional-health-bfbe2483-5108-4f6f-b5b7-b3545de7e2eb?ashby_jid=bfbe2483-5108-4f6f-b5b7-b3545de7e2eb&utm_source=sponsored&utm_medium=newsletter&utm_campaign=q1_oop) - [Agent Engineer](https://www.amigo.ai/careers/agent-engineer-e21088e2-de85-49ce-aea4-cc7228e6b7d9?ashby_jid=e21088e2-de85-49ce-aea4-cc7228e6b7d9&utm_source=sponsored&utm_medium=newsletter&utm_campaign=q1_oop) You can see all the jobs they’re hiring for here: [https://www.amigo.ai/careers](https://www.amigo.ai/careers?utm_source=sponsored&utm_medium=newsletter&utm_campaign=q1_oop) ## **Out-Of-Pocket Take** A few things I think are interesting about Amigo: **Patient-Facing Agents are Here** - It’s clear that we’re entering the takeoff phase for patient facing agents. Amazon launched their AI in One Medical, and several startups like Doctronic and Lotus aim to deliver care directly with agents. Federal tailwinds seem to be pushing for more care delivered via agents - ARPA-H has [a whole program](https://arpa-h.gov/news-and-events/arpa-h-revolutionize-cardiovascular-disease-management-clinical-agentic-ai) for agents to help with cardiovascular issues.  This does feel like the exact right time to have a business around helping providers build and launch their own agentic workflows. Patients are going to eventually start demanding certain workflows be automated as they interact with other services. The stakes are also much higher in these interactions so regulators are going to demand traceability and accountability. Amigo should have the pieces to ride these trends if they execute well. **Building In-House Evaluation + Self Correction Processes** One of the tricky parts of deploying AI is how to evaluate it and then how to correct it (which we’ve [talked](https://www.outofpocket.health/p/how-should-we-evaluate-healthcare-ai-some-thoughts#should-patients-be-bringing-in-more-data) about in the past). Benchmarks published publicly are useful, but can change dramatically when rolled out to a specific setting. It’s clear that as AI tools start having interactions at scale, it’s going to be impossible for humans to investigate all of them. Using synthetic environments and having other AI as the judge/oversight is going to be a necessity (with still some humans in the loop to agree on the changes). Being able to assess the issue and correct the agent is way easier if you control all of the pieces under the hood and understand how each works. If you’re coordinating several different vendors for a task, it’s going to be impossible to understand what went wrong with the other vendors and what fixes should be made. The result is hopefully creating better site-specific and interaction-specific benchmarks to see how an agent would perform in the setting they’ll actually be deployed in. ![Blog image](https://www.amigo.ai/images/blog/content/how-clinical-agents-get-built-and-trained-7.png) **The Benefits of Being Global** - Amigo’s first customer was in Australia, and roughly half of their customers are outside of the US.  I think there’s something interesting about healthcare agents in international markets. - Developing countries outside of the US have a lot more cash-pay volume, so increasing the throughput of interactions with agents has a clear return on investment. - WhatsApp/natural chat interfaces are used way more vs. in the US where you need to go through a patient portal. Those chat interfaces can make healthcare agent interactions more accessible. - Because of how bad accessibility is in many countries, they’re more flexible about deployment of AI tools. There are pros/cons around patient privacy and oversight, but this is happening regardless. - What if agents trained on Johns Hopkins protocols/context graphs can be sold to clinics in other countries? This sort of happens already when Mayo Clinic opens up a hospital in a new country and needs to educate all the local doctors on how they do things; maybe you’ll see something similar to agents. I just hope when agents study abroad they aren’t half as annoying about it as the people who do. As with any company, Amigo may face challenges, and here are some I’d expect as it grows. **Point Solution vs. Platform** A company that picks one specific use case like prior authorization or appointment scheduling can deploy more quickly and prove ROI faster. Amigo requires more upfront work to actually get set up and requires more involvement from the provider to actually build out the agents. Deployment is on average six weeks, but this also buys time as ROI for clinical use cases can take a while.  It’s a bit of a race: can the point solution companies targeting lower-risk tasks get in through the door and then get to the more complex clinical stuff? Or, do you need to come in with the point of view that all agents need to be built on top of the platform in order to do complex clinical tasks? Amigo is betting on the latter. **Will Healthcare Orgs Ever Trust Patient-Facing AI?** Some organizations just aren't ready to let AI talk directly to patients. They’re cool with the back-office, but there’s way more liability when it comes to the front. You think doctors are going to listen to all this mumbo jumbo about world building and simulations that make the agents safe? They’re worried that if something fails, it’s their ass on the line. Amigo is making the bet that the trend is towards more practices willing to actually use agents in patient-facing tasks. **Regulatory Uncertainty** There's ongoing discussion about whether patient-facing AI doing clinical tasks might end up in the software-as-a-medical-device (SaMD) category. The FDA hasn't been super clear here, but in general, this administration seems to be focused on getting AI into clinical practice as quickly as possible. The Doctronic [pilot](https://commerce.utah.gov/2026/01/06/news-release-utah-and-doctronic-announce-groundbreaking-partnership-for-ai-prescription-medication-renewals/) in Utah allowing autonomous prescribing is a litmus test. If this goes well, we’ll probably see more states pushing to allow this kind of AI agent usage. It’ll also put pressure on doctor’s offices to enable this kind of capability.  It’s possible, however, that the FDA decides any clinical, patient-facing AI is considered a medical device, which would significantly increase the burden of proof for agents. ![Blog image](https://www.amigo.ai/images/blog/content/how-clinical-agents-get-built-and-trained-8.png) **Competition** I say this for basically any AI company at this point, but there’s always the question of “what if the EMRs just build this?”. It’s healthcare’s version of “what if Google just copies this?”.  This is always an existential risk. Amigo’s product seems to require stitching together enough other systems that this would be hard for one single system of record to do. On top of that, it’s more likely that the EHRs realize they can actually make money by charging the agents to interact with their systems and enabling that instead. The other vector of competition is foundation models getting into this themselves and helping customers directly deploy agents (e.g. OpenAI [acquiring](https://www.reuters.com/business/openclaw-founder-steinberger-joins-openai-open-source-bot-becomes-foundation-2026-02-15/) the company that made OpenClaw). Will vertical specific training for agents be a huge differentiator? That’s the question for the future. ## **Conclusion and Parting Thoughts** I’ve been building agents a bit myself. They’re extremely powerful, but even for my very simple agents I have to spend a lot of time giving them instructions and rules. I do think we’re going to have more agents in healthcare, but it’s clear that we need a full process to create, shape, and monitor them. Amigo’s bet is that they’ve built a system to do that. But the real agents? The amigos we made along the way. Thinkboi out, Nikhil aka. " Agent Factory? More like Casamigos" Twitter: [@nikillinit](https://twitter.com/nikillinit) IG: [@outofpockethealth](https://www.instagram.com/outofpockethealth/?hl=en) Other posts: [outofpocket.health/posts](https://www.outofpocket.health/posts) *‎If you’re enjoying the newsletter, do me a solid and shoot this over to a friend or healthcare slack channel and tell them to sign up. The line between unemployment and founder of a startup is traction and whether your parents believe you have a job.* --- # CMS ACCESS Model Explained (And How Amigo Helps You Qualify) Source: https://www.amigo.ai/blog/cms-access-model-explained Author: Claire Uhm Published: 2026-02-13 What would your chronic care program look like if Medicare paid you to modernize it? CMS just unveiled the new digital-first **ACCESS model** (Advancing Chronic Care with Effective, Scalable Solutions), a 10-year voluntary program launching July 5, 2026 that shifts Medicare from fee-for-service to **outcome-based reimbursement.** Put simply, **ACCESS will pay providers when patients get healthier**, not when visits or services are billed. In doing so, CMS is giving organizations the flexibility to adopt technology and tools that meaningfully improve patient outcomes. By encouraging technology-supported care, CMS aims to strengthen support for chronic conditions while rewarding clinicians and organizations that deliver real improvements. ## What is the ACCESS model and why did CMS create it? [90% of the $4.9 trillion spent annually on US healthcare](https://www.cms.gov/data-research/statistics-trends-and-reports/national-health-expenditure-data/historical) is for people with chronic and mental health conditions like hypertension, diabetes, kidney disease, depression, and anxiety. These conditions require continuous management, which is fundamentally incompatible with the current visit-based, fee-for-service payment model. CMS created ACCESS to address this problem. Traditional Medicare offers limited reimbursement for the continuous, tech-supported services that help chronically ill patients between appointments. As a result, many patients (especially in rural or underserved areas) lack access to modern tools that could meaningfully improve their health. This new model changes the incentives by tying payments to outcomes rather than activities. Did blood pressure improve? Did pain or mood scores get better? Providers will be encouraged to use innovative digital tools and care models to improve patient health, with the ultimate goal of improving outcomes at scale and reducing long-term costs by keeping people healthier. ## Who is ACCESS for? ACCESS is open to any organization capable of delivering integrated, tech-enabled chronic care. This can include but is not limited to: - Primary care practices - Multispecialty groups and health systems - Orthopedic and musculoskeletal clinics - Women’s health practices - Behavioral health groups - Weight loss clinics - Virtual chronic care companies CMS is explicitly encouraging partnerships between provider groups and tech companies to deliver the full spectrum of care. ## How the ACCESS payment model works At the heart of ACCESS is a new payment model that replaces visit-based billing with recurring payments that reward results. The model includes two components, with Outcome-Aligned Payments for ACCESS participants and an additional Co-Management Payment for non-ACCESS clinicians. ### 1. Outcome-Aligned Payments (OAPs) ACCESS participants receive predictable, recurring payments for managing patients’ chronic conditions. Full payment depends on achieving measurable health improvements across your patient population, such as reducing blood pressure, improving depression scores (PHQ-9), or decreasing pain levels. Organizations earn higher payments when a greater share of patients meet clinical improvement targets for their track, and thresholds rise gradually each year, rewarding sustained success. ### **Payment rates by track** CMS has published the annual OAP allowed amounts per patient for each clinical track. These amounts include both the Medicare program payment (80%) and beneficiary coinsurance (20%), which providers can choose to waive uniformly. ![Blog image](https://www.amigo.ai/images/blog/content/cms-access-model-explained-0.png) The **Initial Period** rate applies when the provider is treating the patient in the clinical track for the first time within the past two years and at least one OAP Measure is not at target. The **Follow-On Period** rate applies for patients who have already been treated in the track or whose health measures are already on target. Providers managing rural eCKM and CKM patients in the Initial Period receive an additional **$15 fixed payment** to offset connected device distribution costs. When a patient is enrolled in multiple tracks with the same provider, CMS applies a **5% discount** to the lowest-cost track(s) during overlapping months. ### **Payment frequency** CMS issues **monthly payments** equal to one-twelfth of the Medicare portion of the annual OAP for valid monthly claims. The sum of monthly payments may not exceed 50% of the Medicare portion of the annual OAP. The remaining 50% is withheld and reconciled after the 12-month care period based on outcome performance. ### 2. Co-Management Payments To support further collaboration, an ACCESS patient’s primary care physician or referring clinician can also bill Medicare for reviewing ACCESS updates and coordinating care. They can receive approximately $30 per review (up to $100/year per patient). These payments have no patient cost-sharing and help ensure PCPs stay actively involved in their patients' chronic care management. ## The four clinical tracks ACCESS is organized into four clinical tracks that group chronic conditions requiring similar longitudinal, tech-enabled management. Organizations may participate in one or multiple tracks. Each track has defined outcome measures that determine whether participants earn Outcome-Aligned Payments. For each measure, a patient meets the target by either reaching an absolute clinical threshold (e.g., systolic BP \< 130 mmHg) or demonstrating a minimum improvement from their baseline (e.g., 15 mmHg reduction) - whichever comes first. ### Early Cardio-Kidney-Metabolic (eCKM) **Conditions covered:** Hypertension *or* any two of the following: dyslipidemia, obesity/overweight with central adiposity, or prediabetes. This track focuses on preventing progression into more serious chronic disease. **Outcome targets:** - **Blood pressure:** systolic BP \< 130 mmHg, or 15 mmHg reduction - **Weight:** BMI \< 30 kg/m² with no more than 5% weight gain, or 5% weight reduction - **HbA1c (prediabetes only):** final HbA1c \< 6.5% - **LDL-C (dyslipidemia only):** final LDL-C \< 100 mg/dL, or 30 mg/dL reduction ### Cardio-Kidney-Metabolic (CKM) **Conditions covered:** Type 2 diabetes, chronic kidney disease (stages 3a/3b), and atherosclerotic cardiovascular disease. **Outcome targets:** - **Blood pressure:** systolic BP \< 130 mmHg, or 15 mmHg reduction - **Weight:** BMI \< 30 kg/m² with no more than 5% weight gain, or 5% weight reduction - **HbA1c (diabetes only):** final HbA1c \< 7.5%, or 1 percentage point reduction - **LDL-C (dyslipidemia/ASCVD):** final LDL-C \< 100 mg/dL (or \< 70 mg/dL for ASCVD), or 30 mg/dL reduction - **eGFR and uACR (diabetes/CKD only):** baseline submission required ### Musculoskeletal (MSK) **Conditions covered:** Chronic pain lasting more than three months, including back pain, osteoarthritis, neuropathic pain, and other conditions that impair physical function and quality of life. **Outcome targets:** Participants select a PROM based on the patient's anatomical site of pain (e.g., PROMIS PF/PI for general pain, ODI for lower back, NDI for neck, QuickDASH for upper limb, KOOS JR for knee, HOOS JR for hip). Additionally: - **Pain intensity (NRS):** no more than 2-point increase from baseline - **PGIC:** end-of-period submission required ### Behavioral Health (BH) **Conditions covered:** Depression and anxiety disorders. **Outcome targets:** - **PHQ-9 (depression):** if baseline ≥ 10, a 5-point reduction; if baseline \< 10, maintain below 10 - **GAD-7 (anxiety):** if baseline ≥ 10, a 4-point reduction; if baseline \< 10, maintain below 10 - **PGIC:** end-of-period submission required - **WHODAS 2.0 (optional):** baseline and end-of-period submission ## How providers are evaluated Two adjustments determine whether participants receive their full withheld payment: 1. **Clinical Outcome Adjustment:** The Outcome Attainment Threshold (OAT) is set at **50%** for the first effective period (July 5, 2026 - December 31, 2027). Participants earn full payment if at least half of their eligible patients meet all required OAP Measure targets. 2. **Substitute Spend Adjustment:** The Substitute Spend Threshold (SST) is set at **90%**. At least 90% of eligible patients must not have received defined substitute services from other Medicare providers for the same condition during their care period. ## Overview of ACCESS requirements At a high level, ACCESS participants must have (or build) the following capabilities: - **Medicare enrollment & clinical oversight:** You must be enrolled in Medicare Part B and designate a physician Clinical Director (MD/DO) responsible for quality and patient safety. - **Data & measurement capabilities:** You need reliable ways to collect patient data, pull in readings from wearables and medical devices, and submit clean digital documentation to CMS. - **Continuous patient support:** ACCESS expects you to check in with patients between visits to collect PROs, monitor progress, answer questions, and keep them engaged over time. - **Outcome tracking & reporting:** You must capture baseline measures, track follow-up results on schedule, and share accurate and complete data with PCPs and other clinicians. ## How Amigo supports ACCESS participation Transitioning to an outcome-based model like ACCESS can reveal gaps in an organization’s capabilities. To participate successfully, organizations will need to collect PROs at scale, track outcomes consistently, engage patients between visits, integrate device data, and report structured metrics back to CMS. Amigo fills these gaps by functioning as the ACCESS operating layer, helping organizations build the capabilities needed to meet program requirements. 1. **Continuous outcomes tracking:** ACCESS pays for outcomes, which requires tracking and improving clinical outcomes like BP, A1c, weight, and PHQ-9/GAD-7 scores. Amigo can handle this entire lifecycle, including automated outreach to collect measurements, structured PRO delivery and scoring, real-time trend detection and escalation when metrics decline, and CMS-ready documentation that flows seamlessly into your EHR. 2. **Automated PRO collection and scoring:** Many organizations struggle to regularly reach out to patients or to administer surveys like PHQ-9 or pain scales systematically. Amigo makes this process effortless by automatically sending, scoring, filing, and summarizing PROs between visits to ensure consistent data without adding burden on human staff. 3. **Continuous chronic care support:** Most chronic care complications happen in the time between visits, which is why ACCESS requires protocols for ongoing care. With Amigo, your practice can deliver continuous care through a clinical expert AI agent that can engage patients 24/7 to send reminders, collect daily vitals, answer common questions, and nudge them toward healthier behaviors. 4. **Multi-condition management:** Because each ACCESS track spans multiple chronic conditions, organizations will need unified workflows that provide a full picture of each patient. A single Amigo agent can manage multiple conditions for the same patient, with full context and tailored follow-up. 5. **Automated documentation & reporting:** ACCESS requires rigorous data tracking that Amigo can automate end-to-end, including vitals, labs, PROs, and wearable/device data such as glucose readings or step counts. All data is captured and logged automatically, then assembled into structured documentation that aligns with CMS reporting requirements. ## Get ready for ACCESS with Amigo ACCESS is a decade-long opportunity to redesign chronic care around outcomes. For organizations willing to embrace technology, it offers a new way to get rewarded for improving patient lives. **Application deadline:** April 1, 2026 **Program launch:** July 5, 2026 **Program runs until:** June 30, 2036 CMS plans to open future cohorts, but early participants will benefit from earlier reimbursement and a longer runway to improve outcomes over the full 10-year period. We’re here to help. **Book a demo with our team** to walk through your current gaps and learn how Amigo can help you meet ACCESS requirements with confidence. For the full CMS payment and performance targets document, see: [ACCESS Model Payment Amounts and Performance Targets (PDF)](https://www.cms.gov/priorities/innovation/files/access-payments-amts-perf-targets.pdf) For the latest updates, visit the [official CMS ACCESS page](https://www.cms.gov/priorities/innovation/innovation-models/access). --- # The Third Leading Cause of Death is Preventable Source: https://www.amigo.ai/blog/actions-architecture Author: Claire Uhm Published: 2026-01-05 How Amigo's Actions architecture enables AI agents to safely and reliably execute clinical tasks. *This is part four in a five-part series diving into the Amigo Cognitive Architecture. We've already covered* [*Functional Memory*](https://www.amigo.ai/blog/functional-memory)*,* [*Context Graphs*](https://www.amigo.ai/blog/context-graphs)*, and* [*Dynamic Behaviors*](https://www.amigo.ai/blog/dynamic-behaviors)*. In our final post, we'll explore the Agent Core.* Medical errors are the [third leading cause of death in the United States](https://www.ncbi.nlm.nih.gov/books/NBK499956/). It's a troubling statistic, but not a surprising one when you consider the interdependent nature of the healthcare system. Healthcare is built upon millions of everyday actions that depend on long chains of people, systems, and handoffs. In an environment this complex, mistakes in execution are inevitable. As clinical AI agents take on more consequential roles in delivering care, they inherit the same underlying reality. This raises an uncomfortable question for anyone building or evaluating healthcare AI: if clinical agents will inevitably make mistakes, how do we make them safe? ## Smarter Doesn’t Mean Safer The intuitive answer would be to simply make the agents smarter, but *reasoning* and *executing* are different problems. More advanced models can help agents make better decisions, but intelligence alone can’t prevent errors in execution. This is also true for human clinicians. A seasoned physician might know exactly which lab tests a patient needs and still occasionally input the wrong code in the EHR. Instead of relying on an agent to take on tasks with unbounded complexity at a zero-percent error rate, the safest approach is to design systems that: 1. Outline precisely how each type of action should be performed 2. Include safety nets to minimize harm should an error occur In healthcare, this is the difference between a recoverable mistake and a preventable death. Our thesis is simple: in order for clinical agents to perform actions safely, every workflow must be **purpose-built** and **safe to fail**. This is the principle behind **Amigo Actions**, an execution layer within Amigo’s cognitive architecture that ensures all clinical workflows run in a controlled environment, creating a safer way for our agents to perform patient-facing work. ## How Do Agents Perform Actions? An AI agent is a software program that can think, make decisions, and perform tasks. In a clinical setting, an agent might retrieve lab results, provide care instructions, or send follow-up reminders. To complete these tasks, agents call on tools (APIs), which are pieces of code that carry out specific actions. Agents are being used to solve increasingly complex problems across industries, resulting in a need to integrate with a progressively larger set of tools over time. In response to this, [Model Context Protocol (MCP)](https://modelcontextprotocol.io/docs/getting-started/intro) was introduced as a universal “tool language” that all agents can speak, allowing them to walk up to any compatible tool provider and ask for different services. MCP is useful as a communication protocol that standardizes *how an agent can ask for information*, but it doesn’t govern *how that request is carried out*. There’s nothing that specifies where the code runs, who sees the data, or what happens if something fails. Most off-the-shelf agent solutions focus on connecting agents to tools via MCP. This is a reasonable approach for low-stakes use cases where a failed action results in a bit of wasted time, but healthcare is different. Even a partial failure can have a direct human cost. If a clinical agent is given access to a tool to update prescriptions, it may start using it in an inconsistent manner. And worse, if it crashes in the middle of performing the action, the patient could end up with a corrupted file or partial record that omits crucial information. ## Safety Lives in the Substrate That’s why we built **Actions**, an execution layer or *substrate* that governs how tasks are carried out. If MCP is a *language* that agents can use to make requests, a substrate is the *environment* in which those requests get executed. You can think of each Action as a **sterile single-use operating room**. Each Amigo agent is equipped with a unique set of these “operating rooms” that are purpose-built for each type of action it is allowed to perform. They are specifically scoped to handle tasks that need to be deterministic and exact, without the unpredictability of relying on an AI model to figure it out. After each operation, the operating room is reset completely. This design ensures that if something goes awry mid-procedure, the impact is self-contained and can be reversed or corrected without affecting the rest of the hospital. This prevents partial failures, acting as an essential failsafe for healthcare applications. By contrast, a typical MCP-based setup relies on whatever environment the agent happens to be running in. There’s no guarantee of how or where an action is carried out, who has access to the data, or if the results can be undone. It’s the difference between a high-tech surgical suite and a hospital with no oversight or standards for keeping records. In one, all of the tools are neatly laid out and your chart is meticulously kept up to date. In the other, the doctor has already forgotten your name and is bringing in the next patient before they’ve had a chance to wash their hands. ## The Anatomy of an Action We designed Actions to be **atomic**, **ephemeral** (or temporary), **auditable**, **scalable**, and **deterministic**. In a nutshell, this means that each time an agent needs to perform a task, we spin up a fresh, isolated workspace with only the resources and permissions it needs. Once the task is done, the workspace self-destructs and any sensitive data is scrubbed clean. ![Blog image](https://www.amigo.ai/images/blog/content/actions-architecture-0.png) Each of these properties play a specific role in keeping agent operations safe and predictable. - **Atomic:** Each Action either *completes fully* or *fails cleanly* so no partial or half-finished task can corrupt patient records or leave clinical workflows incomplete. - **Ephemeral:** Every Action runs in an isolated environment that gets torn down and wiped clean after use, so no sensitive data persists once the task is finished. - **Auditable**: Every Action is logged, generating a paper trail of what was done, when, and by which agent. This visibility allows teams to review activity, trace errors, and roll back to previous versions when needed. - **Scalable:** Actions can expand from a single request to thousands running in parallel, all governed by the same substrate. This ensures consistent performance and predictable costs, even during peak demand. - **Deterministic:** Each Action runs a pre-defined workflow with only the resources and permissions required, producing consistent outcomes. ## How Amigo Actions Make Clinical Agents Safer To appreciate the difference that Actions make in practice, consider the following scenario with an agent that processes medication refill requests. ### **Preventing Errors** In the absence of a safety substrate like Actions, the agent might have broad access to the prescribing system and figure out how to complete the task on its own, which introduces room for unpredictable behavior and outcomes. With Actions, nothing is left to the model's judgment. The refill workflow is precisely defined in advance, including what data to pull, which checks to run, and in what order. In this example, the Action might include checks to cross-reference the patient's chart and verify that the medication is appropriate for their recorded conditions and allergies before placing the order. If anything is surfaced, the process can be terminated, the care team notified, and the incident logged. ### **Containing Errors** What if the agent attempts to place the order but the process crashes halfway through? Under a typical setup, this could result in a partial order where the prescription is logged in the patient's chart but never sent to the pharmacy. With Actions, each operation is atomic, so the medication order completes fully or not at all, and there's no scenario where a partial order goes through. ### **Recovering From Errors** Once an error is contained, teams can quickly identify which version of the Action was at fault, roll back to a safe version within seconds, and even use the audit trail to cancel or adjust any orders that were placed incorrectly. The Actions architecture creates a clinical-grade safety net, allowing healthcare organizations to deploy patient-facing agents with confidence. ## Trusted Execution for Safe, Scalable Agents The safest systems in healthcare work on two fronts: minimizing errors and limiting the blast radius when they occur. When it comes to clinical agents, safe execution is non-negotiable and our Actions architecture makes it possible at scale. We’re building the frontier of safe and intelligent clinical agents - the kind you can trust with lives. Ready to see what clinical-grade AI looks like in practice? Book a call with us today. --- # Why Most Healthcare AI Can't Handle Real Patient Conversations Source: https://www.amigo.ai/blog/dynamic-behaviors Author: Claire Uhm Published: 2025-12-17 Amigo's Dynamic Behaviors allow AI agents to recognize what matters during clinical conversations and adapt in real time. *This is part three in a five-part series diving into the Amigo Cognitive Architecture. We've already covered* [*Functional Memory*](https://www.amigo.ai/blog/functional-memory) *and* [*Context Graphs*](https://www.amigo.ai/blog/context-graphs)*, and in upcoming posts, we will explore Actions and the Agent Core.* ## Why Can't Healthcare Agents Think Like Doctors? Consider the following exchange between a patient and an AI intake agent. The patient says: > My chest feels a bit tight. The AI agent replies: > I’m sorry to hear that—chest tightness can be commonly caused by anxiety, acid reflux, or muscle strain. Let’s continue with your intake. What is your date of birth? A doctor would never respond this way. A good clinician listens to patients with a running list of hypotheses in mind, constantly reassessing new information to determine what matters and what the next step should be. When did the symptoms start? Has the patient started taking any new medications recently? Does their medical history make anything more likely? The conversation itself is diagnostic and each answer reshapes the questions that follow. We call this *clinical judgment* - the ability to recognize when something is important and respond accordingly. Most AI agents in healthcare today can’t do this. Some are built to execute a set of predefined workflows very well but aren’t able to change course when new information surfaces. Other agents are able to have more flexible conversations but lack the infrastructure to process new context in a structured way. As a result, they struggle to ask the right follow-up questions and determine the best path forward in real time. **Dynamic Behaviors** are what allow Amigo agents to exercise clinical judgment like human clinicians. But building this kind of adaptability requires solving a fundamental problem in how AI models access memory and apply knowledge. Traditional models often struggle here due to what we call the *latent space activation challenge*. ![Blog image](https://www.amigo.ai/images/blog/content/dynamic-behaviors-0.png) ### The Latent Space Activation Challenge AI models can possess the right clinical knowledge and still fail to apply it correctly. Inside every large language model (LLM) is a *latent space*, or a massive library where all its knowledge is stored. The library is vast, and success at knowledge retrieval depends on whether the model can locate and combine the right pieces of information for the problem at hand. When this happens without strong control architecture in place, the model is simply guessing. It may solve the wrong problem or prioritize incorrectly, often producing surface-level outputs that look right but lack real understanding. In healthcare, this can be dangerous. Consider the patient experiencing chest pain. It could be caused by acid reflux. It could also be a sign of a heart attack. A physician knows this, but that knowledge is only useful if they're able to ask the right questions. AI models face the same issue. All the foundational medical knowledge is there, but accessing it incorrectly produces inconsistent, irrelevant, and unsafe results. Dynamic Behaviors are an architectural innovation that solves this problem by ensuring that Amigo agents activate the right clinical knowledge at the right moment**.** ### How AI Agents Navigate Complex Conversations In our previous Deep Dive, we discussed how [Context Graphs](https://www.amigo.ai/blog/context-graphs) give Amigo agents a structured yet flexible roadmap for clinical conversations. You can think of Context Graphs as the scaffolding that breaks down complex clinical workflows into connected states, giving agents reliable options to intelligently navigate through intake questions, treatment discussions, or care planning. This solves the problem of agents getting lost or going completely off-script. But structure alone isn’t enough. Clinical conversations can be unpredictable, and patients will sometimes have unusual queries. An agent might be halfway through a medication review when the patient mentions something unexpected or concerning. Dynamic Behaviors allow Amigo agents to recognize these critical moments and override the conversation to deliver nuanced responses without abandoning the underlying structure underneath. This keeps patient interactions flexible and natural, instead of feeling like rigid, predetermined conversational pathways. ## How Dynamic Behaviors Work From an architecture perspective, Dynamic Behaviors act as a *trigger-and-response* mechanism relying on two components: triggers and instructions. **Triggers** detect patterns that signal when a behavior should activate. These can range from explicit keywords (“I have chest pain”) to subtle contextual cues (a patient repeatedly deflecting questions about medication adherence). Amigo’s trigger system doesn’t rely on a single signal to decide when a behavior should activate. Instead, it draws on multiple cues, such as the agent’s internal reasoning, past responses, patient context, and tool usage. When these signals combine, each new behavior is informed by the full history of the patient relationship rather than just the last message. As a result, Amigo agents are able to recognize patterns that would be invisible to simpler systems. To illustrate, consider a patient who consistently reports low energy, sleep disruption, and loss of interest in activities they previously enjoyed. Even if they don’t explicitly report feeling depressed, a trigger can activate to prompt the agent to explore this possibility further. **Instructions** define how the agent responds once triggered. They can be broad, giving the agent discretion to adapt, or highly detailed, requiring strict adherence to a defined protocol. A safety escalation might have rigid instructions, such as immediately advising the patient to call 911, whereas a trigger for exploring lifestyle factors might give the agent more flexibility in how it approaches the conversation. These instructions can also tell the agent to perform external actions, like placing a call to a specialist’s office or writing to an EHR. ![Blog image](https://www.amigo.ai/images/blog/content/dynamic-behaviors-1.png) ## Why Dynamic Behaviors Matter for Healthcare AI Dynamic Behaviors change what’s possible in clinical AI. 1. **Increased safety**: Rather than relying on patients to explicitly state their concerns, Amigo agents can detect risk based on implied patterns or contextual triggers, then control the response based on custom-defined safety protocols. 2. **Improved patient experience**: Conversations feel natural rather than scripted because Amigo agents respond to what patients are actually saying rather than what the workflow anticipated. Patients rarely take the happy path, and a robust AI experience must be built for edge cases. 3. **Execute complex workflows:** Dynamic Behaviors can trigger tools, connect with external systems, or pull live data mid-conversation. ## Deliver Responsive, High-Quality Care with Amigo Dynamic Behaviors allow Amigo agents to recognize critical moments in clinical conversations and respond appropriately, whether that means asking different questions or escalating to a clinician. The result is healthcare AI that can deliver responsive, high-quality care at scale. Ready to build healthcare agents that do more than follow a script? Book a call with us today. --- # Amigo Establishes Medical Advisory Board and Appoints Dr. Jay Shah as Chief Medical Advisor Source: https://www.amigo.ai/blog/amigo-establishes-medical-advisory-board Author: Richard Wang Published: 2025-10-30 Stanford Health Care's Chief of Medical Staff to Guide Clinical Strategy for Healthcare AI Platform *Click *[*here*](https://www.prnewswire.com/news-releases/amigo-establishes-medical-advisory-board-and-appoints-dr-jay-shah-as-chief-medical-advisor-302599081.html)* to read the following press release on PR Newswire.* --- NEW YORK, Oct. 30, 2025 /PRNewswire/ -- Amigo, an AI platform that partners with healthcare organizations to build safe and reliable clinical agents, today announced the formation of its Medical Advisory Board with the appointment of Dr. Jay Shah, MD as Chief Medical Advisor. Dr. Shah, who currently serves as Chief of the Medical Staff at Stanford Health Care, will guide Amigo's clinical strategy and ensure it meets the stringent standards of healthcare delivery. The Medical Advisory Board represents Amigo's commitment to building agentic solutions grounded in clinical expertise and responsible AI practices. As the board's inaugural member, Dr. Shah will work directly with Amigo's leadership team to shape product development, establish clinical validation frameworks, and ensure the platform addresses real-world healthcare challenges. "To responsibly deploy AI in healthcare you need to understand the complexity of clinical care, not just the technology," said Ali Khokhar, CEO at Amigo. "Because we are pioneering a new model of care, it's essential that we involve physicians in designing this new approach. Dr. Shah brings exactly the perspective we need, and his guidance will be invaluable as we help healthcare providers responsibly scale their clinical expertise." Dr. Shah is a cancer surgeon and associate professor of Urology at the Stanford University School of Medicine. He is a nationally recognized expert in robotic surgery and bladder cancer treatment. Prior to Stanford, Dr. Shah served as Center Medical Director for the Genitourinary Center at MD Anderson Cancer Center, where he launched the bladder cancer robotics program and developed an enhanced recovery program for patients undergoing bladder removal surgery. He is a graduate of Harvard College and completed his medical degree and residency at Columbia University, where he was elected to the Alpha Omega Alpha Medical Honor Society, named Physician of the Year, and recognized with the Gold Foundation Excellence in Teaching Award. "AI has incredible potential to improve patient outcomes, but only if it's built with deep clinical understanding and proper safeguards," said Dr. Shah. "I'm joining Amigo because I believe what they're building will transform care delivery, and that they are building it with transparency and clinical rigor." Amigo's Medical Advisory Board will bring together a thoughtfully assembled team of clinical leaders from multiple disciplines to guide the company's product roadmap and ensure it meets healthcare's unique regulatory and operational requirements. **About Amigo** Amigo builds control and reliability infrastructure to enable the development of secure, compliant, and clinically validated AI agents. The company partners with healthcare providers to build clinical, patient-facing agents that scale patient access and deliver quality care at the cost of compute. Amigo's mission is to enable the transition from AI as experimental technology to trusted infrastructure for the economy's most critical functions. For more information, visit [**Amigo's**](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4544694-1&h=2339960405&u=https%3A%2F%2Fwww.amigo.ai%2F&a=Amigo's) website. **Media Contact:** Richard Wang richard@amigo.ai SOURCE Amigo Inc --- # The path to patient-facing AI? Follow Waymo's lead Source: https://www.amigo.ai/blog/the-path-to-patient-facing-ai Author: Christina Farr & Anjalee Khemlani Published: 2025-10-29 5 principles for patient-facing AI, according to top clinicians *The following is a direct transcription of the article. To read this post on Second Opinion's website, please visit* [***this link***](https://secondopinion.media/p/the-path-to-patient-facing-ai-follow-waymo-s-lead)*.* --- How would you feel about a doctor telling you, “Hold on, let me confirm with ChatGPT about that”? It's already happening among health professionals via tools including OpenEvidence, Doximity and UpToDate, as well as (informally) ChatGPT. And so is the reverse: Patients are using ChatGPT to ask questions about their symptoms or disease before talking to a doctor, as well as analysis related to their lab work or imaging. The approach marks a shift in a world where physicians rolled their eyes when Dr. Google walked in the room, knowing the information could be wrong. At an event recently for clinicians in digital health, one panelist commented that it used to be considered lazy when a physician used UpToDate to look something up. Now, it’s considered lazy not to leverage AI. Dr. David Rhew, Global Chief Medical Officer at Microsoft, said that AI is forcing important questions about which doctors should be leveraging AI, and for what purposes. Doctors can make mistakes, but it’s early days for the technology – which remains prone to errors and fabrications. “We tend to forget that doctors are human. There is a huge variability in how doctors care for patients,” Rhew told *Second Opinion.* Now, as agentic AI tool pitches are finding their way into every corner of healthcare, from back end support and administrative tasks through patient care and maintenance, experts are thinking about how to approach patient-facing tools. And in its current form, the industry lacks a few key foundational needs: trust, a widely accepted framework, and training for all stakeholders. Recently, [Stanford and Harvard Universities teamed up](https://www.forbes.com/sites/saibala/2025/10/07/meet-arise-a-stanford-and-harvard-backed-lab-dedicated-to-objectively-validating-ai-in-healthcare/?utm_campaign=the-path-to-patient-facing-ai-follow-waymo-s-lead&utm_medium=referral&utm_source=secondopinion.media) to help overcome one barrier: validating AI tools in healthcare. But there’s still a ways to go. “We’re out of early adolescence right now, with these tools, maybe even earlier than that,” said Harlan Krumholz, Director of the Yale New Haven Hospital Center for Outcomes Research and Evaluation. Krumholz noted that there’s still an adoption curve around how to correctly use the tools. “They’re kind of being used in the way you would use a search, but they are better than a search engine.” It’s also not necessarily the case yet that humans are being “augmented” by AI and that the combination of the two is the most accurate and effective. [Studies](https://www.nature.com/articles/s41562-024-02024-1?utm_campaign=the-path-to-patient-facing-ai-follow-waymo-s-lead&utm_medium=referral&utm_source=secondopinion.media) have found that communication barriers, trust issues, ethical concerns and a lack of good coordination can play a role in hindering this collaboration. Although that’s starting to resolve as newer studies are finding in some [specialties](https://journals.lww.com/ajsp/fulltext/2018/12000/impact_of_deep_learning_assistance_on_the.7.aspx?utm_campaign=the-path-to-patient-facing-ai-follow-waymo-s-lead&utm_medium=referral&utm_source=secondopinion.media), but not all, AI plus human is the best combination. Another [recent study showed](https://www.michiganmedicine.org/health-lab/adults-dont-trust-health-care-use-ai-responsibly-and-without-harm?utm_campaign=the-path-to-patient-facing-ai-follow-waymo-s-lead&utm_medium=referral&utm_source=secondopinion.media) that 66% of adults have low trust in their health systems to use AI responsibly, and more than half said they don’t trust health systems to use an AI tool that would not harm them. So how can we overcome this trust gap – as a society and within the industry? Second Opinion spoke to several other experts about what 5 guiding principles of patient-facing agentic AI looks like. Here’s what the experts had to say. ## Guiding principles ### **Build trust** The trust gap in healthcare for use of AI exists both on the patient side as well as the clinical side. And there is only one shot the industry has to get it right – because once trust is lost in the technology, it will be a hard fight to get it back. Especially with an already-dubious clinician population. [Amigo.ai](https://amigo.ai/?utm_campaign=the-path-to-patient-facing-ai-follow-waymo-s-lead&utm_medium=referral&utm_source=secondopinion.media) CEO Ali Khokhar said companies have to do for healthcare what Waymo has done for autonomous driving. What was so effective about Waymo’s strategy wasn’t just that it was transparent. It was the incredibly careful approach the company took, including in countless simulated performances before the autonomous vehicle hit the streets. In healthcare, building that trust will take time - and a lot of research. Also along these lines, it was a physician – Dr. Jonathan Slotkin – who went [viral](https://www.linkedin.com/posts/slotkinjr_as-a-neurosurgeon-i-care-a-lot-about-road-activity-7374146589012942848-CUKV/?utm_campaign=the-path-to-patient-facing-ai-follow-waymo-s-lead&utm_medium=referral&utm_source=secondopinion.media) in the past few weeks for digging into the data and determining that Waymo is doing more than an incremental improvement in safety. “It’s categorical.” In the same vein, physicians will adapt when they’re presented with compelling enough data. It’s a role that is grounded in the scientific method, but adaptation is also necessary in the face of massive physician shortages. “I think AI doctors are more similar to self-driving cars than any other analogy I could give you,” said Khokhar, noting that he started his own company to bring Waymo-like thinking to healthcare. “I think of self-driving cars as the only place where...the cost of failure is somebody can die,” Khokhar said. ### **Specialty-based training** Each tool needs to be able to address the variations in disease-specific and specialty-specific care. Having a one-size-fits-all approach is going to relegate any tool to the equivalent of a search engine, rather than the power of AI providing augmented support for doctors. Again, similarly to Waymo, that means endless hours of virtual simulation training, and ensuring there are no hallucinations or biases. Some experts worried the tools could reverse progress made in the real world, and amplify existing biases in medical care – which has taken the industry decades to only just begin to address in recent years. “Medicine is deeply contextual. What’s relevant in radiology looks very different from what’s needed in cardiology,” said GE HealthCare Global Chief of Science and Technology Officer Taha Kass-Hout. He believes the long-term vision for the industry is to have a multi-agentic AI platform where the specialty-trained agents can collaborate to ensure the most valuable and appropriate clinical care is given. ### **Include physicians** Not just in the build-out of tools, but also in the ability for patients and physicians to be connected to the information being sought and shared. This can solve one of the largest gaps in digital health to-date: leaving the tools outside the exam room. The solutions should include integration with digital health records, which gives doctors visibility into the patients’ needs. And it helps the doctor stay actively involved with their patients’ care, rather than the current route of falling off the radar once they exit the clinic. Khokhar from [Amigo.ai](https://amigo.ai/?utm_campaign=the-path-to-patient-facing-ai-follow-waymo-s-lead&utm_medium=referral&utm_source=secondopinion.media) stressed that close integration and collaboration with physicians and their existing workflows is key. Manmeet Kaur, Executive in Residence at the Regenstrief Institute and a Healthcare Entrepreneur, said integration is everything. “A lot of the patient monitoring wave of digital technology limitations we’re facing is not being integrated with the physician’s care,” she explained. And without that – “It can become noise, it can become less relevant,” Kaur said. ### **Training for all stakeholders** Ensure that whoever is going to be using the tools, whether it be a patient or a health professional, is trained fully on the limitations and best ways to use the tools. This can prevent harm and also ensure that even hospital executives and other figureheads in the system who are not in touch with patients also understand how this tech is being used. There are some hospitals in other parts of the world that are building AI at the center of their health system rather than just tacking it on, and that requires a system-wide training. Similarly, health execs and professionals need to be thinking about how to integrate the product and maximize the information database. Will Morris, a Physician and Executive, said it’s key that these tools are used to help inform patients, and help them make decisions. But not to be used for self-diagnosis. “Let’s not forget what we had before…the interweb. We had missing information. You had information asymmetry.” When patients are empowered, and doctors are helped, with additional information at their fingertips, it can help strengthen the relationship and treatment process, Morris said. That can only become reality if the tools are trained on authoritative sources and provide the best information to all sides without hallucinations and biases, he explained. ### **Regulations, malpractice, liability** There are mixed thoughts about how to create an industry-wide framework that helps regulate and monitor the build out of these tools. And should there be a validating body of sorts? And what does malpractice insurance and liability look like once medical professionals begin to rely on these tools? Those are questions that need answering. Some experts want a validating body, while others believe that it would create more red tape and slow down innovation, and possibly create more barriers. Ultimately, the experts all agreed that there’s no going back. Stephen Klasko, Executive in Residence at General Catalyst and a former Health System Executive, told us that consumers have reached a “breaking point” with healthcare’s fragmentation. According to Klasko, AI can play a role in integrating the care they receive within the four walls of a health system, and outside of it: “The pandemic accelerated digital health adoption,” he said. “But also exposed the chaos of all these Lego pieces, these digital health tools.” *Are we missing any key guiding principles that should govern the use of patient-facing AI? Reach out to us and let us know. We’d love to hear from you!* --- # Deep Dive: Context Graphs Source: https://www.amigo.ai/blog/context-graphs Author: Ali Khokhar Published: 2025-10-08 How Amigo's Context Graph architecture enables structured yet flexible clinical conversations. *This is part two in a five-part series diving into the Amigo Cognitive Architecture. We've already covered *[*Functional Memory*](https://www.amigo.ai/blog/functional-memory)* – in upcoming posts we will explore Dynamic Behaviors, Actions, and the Agent Core.* Almost all AI agents fall short for real-world healthcare use cases because they can't balance structure with flexibility. They're either too rigid to handle the nuance of real patient cases, or too loose to maintain the clinical rigor that safety demands. Amigo solves this problem with context graphs – a proprietary architecture that reimagines how to navigate complex conversations. Unlike traditional agent frameworks that rely on either linear decision trees or largely unstructured reasoning, context graphs provide external scaffolding that preserves clinical logic while enabling personalized adaptation. They lay out the broader landscape of a task, its purpose, its structure, and the intricate connections that hold it together. The outcome is agents that can follow vetted protocols as reliably as the best clinicians, while flexing to each patient's unique circumstances. *How do you teach an AI agent to navigate complex clinical conversations at scale?* This was the core question that led us to design the context graph, a blueprint that helps agents navigate a specific problem space by breaking it down into smaller, connected parts. Picture a climber scaling a mountain face. The context graph represents pre-mapped footholds that show the climber (the agent) validated pathways to navigate upward. These footholds present them with multiple safe paths they can choose based on real-time factors like weather conditions or available equipment. ![Context graphs enable situational state navigation without the rigidity of standard agent flowcharts or decision trees.](https://images.amigo.ai/blog/content/context-graphs-0.webp) ### Balancing Structure with Flexibility Having a clear structure is crucial – a PCP seeing a patient for their annual visit follows a specific flow that is comprehensive and likely to surface critical issues. But cases are often complex and every patient requires a degree of personalization based on factors like medical history and health literacy. This interplay means the agent must appropriately balance clinically tried-and-true structure with the ability to adapt as needed. Context graphs solve this by offering a flexible spectrum that clinical teams can calibrate to their specific requirements: - **Strict contexts** for critical clinical protocols (e.g., medication instructions, safety procedures) - **Medium-flexibility contexts** for clinical guidance (e.g., treatment discussions, care planning) - **Open-ended contexts** for patient conversations (e.g., empathetic guidance, building rapport) ### Bridging the Token Bottleneck Another critical function of the context graph architecture is that it overcomes a significant limitation faced by current AI models that we call the *token bottleneck*. Imagine a genius with short-term memory loss who reasons by writing down one word at a time. Each time they write a word, they completely lose their memory and have to reconstruct their reasoning by reading previous words. This is the challenge faced by today's LLMs. When models need to "think through" problems step-by-step, they must compress their reasoning into text tokens to express it. One token is emitted and the internal state is effectively reset. The model must then rebuild context from its output. This causes a significant loss in reasoning that is unacceptable in a high-stakes, high-complexity domain like healthcare. Clinical decision-making involves vast amounts of interconnected information like patient histories, regulatory requirements, and clinical pathways, yet foundation models cannot hold onto the majority of this context. We designed context graphs to act as external scaffolding for agents to organize and preserve their reasoning. Instead of letting the model lose track of the conversation as it progresses, the context graph holds important details in place so the agent can frame its responses correctly. For example, an intake agent talking to a patient will know at all times what they've already covered and where the conversation needs to go next, allowing it to stay on topic and frame its questions appropriately. ### Anatomy of a Context Graph At its core, a context graph organizes decision points into layered states, enabling an agent to navigate complex behaviors efficiently. Instead of branching in one direction along preset "rails" like a decision tree, a context graph works more like a spiderweb; the agent can travel forward, backward, or laterally along a number of valid pathways based on situational context. There are multiple types of states, each playing a distinct role in guiding an agent's behavior and managing conversation flow: - **Action states** execute tasks or respond to the user within established rules and constraints, guided by the active conversational context - **Decision states** determine optimal actions based on real-time data and objectives, simultaneously drawing on memory, knowledge, and reasoning - **Reflection states** enable deeper thought and force the agent to carefully consider its reasoning before moving forward - **Recall states** explicitly retrieve user memory or past interactions to personalize and improve responses, bringing historical context into play - **Annotation states** clarify and segment complex interactions to help the agent keep tabs on key information - **Side-effect states** identify points where the agent can interact with external systems or trigger actions outside of its own environment Amigo's Agent Engineers work with our partners' clinical teams to deeply understand the structural topology of the problem the agent is meant to solve, then use these core building blocks to construct the optimal context graph. This state-based architecture also allows the agent's reasoning to be broken down into clear, traceable steps that make it possible to audit conversations with extreme visibility. ### Layering in Functional Memory Context graphs and [functional memory](https://www.amigo.ai/blog/functional-memory) work as complementary systems. While the context graph provides the structural roadmap for clinical conversations, the memory system ensures the agent navigates that roadmap with the right patient knowledge at the right time. The agent's user model stays active throughout navigation, providing continuous access to the complete patient picture as it moves through the context graph. This enables the agent to make informed decisions at each junction and respond appropriately based on the patient's clinical profile. Depending on what the context graph dictates, the agent can also dig deep to perform *memory expansion* to retrieve specific information missing from the high-level user model. This can happen either implicitly (when the agent determines it has insufficient information to execute on a task demanded by the context graph) or explicitly through a recall state that forces deliberate historical recontextualization. Together, these systems create a feedback loop: functional memory ensures the agent navigates the context graph with the right contextual framing, which in turn helps the agent access and update its memory more effectively. ### Structured Intelligence for Personalized Care Like a seasoned clinician who follows established protocols while adapting to each patient's unique circumstances, Amigo's context graph architecture provides the structural foundation for intelligent clinical reasoning. By organizing complex healthcare conversations into interconnected states rather than rigid decision trees, these graphs enable agents to maintain clinical rigor while preserving the flexibility essential for personalized patient care. Working seamlessly with the functional memory system along with the rest of the Amigo Cognitive Architecture, context graphs ensure that every interaction is both clinically sound and contextually appropriate, allowing healthcare organizations to scale expert-level services across each patient while maintaining the nuanced decision-making that defines quality care. Interested in how Amigo applies context graphs to build more intelligent clinical agents? [Book a call](https://cal.com/richardwang/discovery-call) with us today. --- # Amigo Deep Dive: Digital Health Wire Source: https://www.amigo.ai/blog/digital-health-wire Author: Jason Barry Published: 2025-09-23 AI moves fast, but trust moves slow. That's why Digital Health Wire is launching a new series to spotlight the companies taking AI from promise to practice. *The following is a direct transcription of the article. To read this post on Digital Health Wire's website, please visit *[*this link*](https://digitalhealthwire.com/co-creating-confidence-inside-amigos-approach-to-building-trustworthy-ai-agents/)*.* --- # Co-Creating Confidence: Inside Amigo’s Approach to Building Trustworthy AI Agents AI moves fast, but trust moves slow. That’s why Digital Health Wire is launching a new series to spotlight the companies taking AI from promise to practice. **First up: Amigo.** No matter how many medical licensing exams and curated case vignettes the latest models conquer, they’ll still need to make it through the proving ground of real clinical practice to get doctors on board. The biggest challenge for AI in healthcare isn’t building agents that can handle a task, it’s building agents that clinicians can trust to handle those tasks safely – every time, guaranteed. There’s a massive gap between textbook performance and real-world reliability, and Amigo is giving providers the infrastructure to bridge that gap. **Earning trust takes more than technology. **Amigo’s process is just as important as its platform for enabling healthcare orgs to safely design, test, and monitor agents that they can genuinely depend on for their unique clinical and administrative workflows. Amigo’s approach to building trust stands on four core pillars: - Controllability – Clinical teams can define and adjust agent behavior. - Performance Validation – High-fidelity patient simulations stress-test readiness. - Real-time Observability – There’s full transparency into decision-making. - Continuous Alignment – Agents adapt to changing protocols and priorities. **“Good enough” isn’t enough in healthcare.** Most industries can get away with using the 80/20 rule to fine-tune their products. If they can improve the experience for 80% of their users, it justifies any shortcomings for the other 20%. Traditional benchmarks might work for customer service, but not when that 20% includes life or death situations. - When AI developers chase benchmark scores but ignore outcomes, they miss the actual point of care delivery: making patients healthier. A perfect medical licensing exam is great, but it’s not the same thing as a perfect clinician – or a trustworthy AI agent. - Strong benchmark scores can also lure providers into a false sense of security, and it’s tough to notice when performance starts to drift if nobody is on the lookout. **Drift is inevitable, and the current is strong.** Even if an AI agent works on day one, there will always be a tendency for performance to slip over time. Clinical guidelines change. New drugs enter the market. Populations evolve.  Amigo safeguards against this drift with a three-layer framework: - The Problem Model – Customers define their specific needs and the “operable neighborhood,” which is basically the set of scenarios that the agent can help with. - The Judge – Customers establish their own success criteria, as well as the verification measures to keep track of them. That includes both safety metrics like accuracy and handoff reliability, plus experience metrics like empathy and response time. - The Agent – Amigo spins up an agent that can safely tackle the problem at hand, then continuously monitors it against the “success scorecard” to minimize drift and intervene well before it impacts patient care. **How can performance be guaranteed?** Simulating success ahead of time. Amigo swaps generic benchmarks for millions of simulated patient conversations to make sure each of its agents are 100% operationally ready before they’re actually deployed. - The simulations reflect the real-world scenarios and demographics of each customer’s unique patient population. The goal is to stress-test the agents to their breaking point in a controlled environment, then refine them until they perform reliably under pressure. - Amigo intentionally oversamples rare scenarios – like patients with unusual drug interactions – to ensure edge cases don’t slip through. This not only helps keep the agents consistent at scale but also means that they frequently perform better in real practice. **It’s a proven blueprint.** Amigo’s strategy for building trust in AI resembles the playbook used in another area with similarly high stakes, high variance, and high skepticism: self-driving cars. - Waymo defines the well-charted terrain where its autonomous vehicles (AVs) are designed to operate safely. Amigo maps specific clinical neighborhoods. - Waymo simulates edge cases that might take years to encounter in the field before its AVs see any actual street time. Amigo puts its agents to the same test. - Waymo’s initial rollout includes safety drivers that can take control when needed. Amigo works with clinicians to refine the accuracy of the Judge. - Waymo removes safety drivers as its AVs prove themselves on real trips. Amigo reduces human oversight once clinicians are confident the Judge is calibrated correctly. - Waymo moves to similar neighborhoods only after success is consistently demonstrated. Amigo can expand to adjacent use cases where its agents can inherit validated behaviors and guardrails. **Adoption follows confidence.** When clinicians co-create the solution to their problems, they’re more comfortable putting it in front of patients.  - That confidence usually means leveraging Amigo to automate the workflows that have been weighing them down the most, such as around-the-clock support and care navigation. - The agents go beyond providing advice. They can perform actions like ordering tests, updating the EHR, and looping in care teams for complex workflows like triage and medication management. **AI still has a lot to prove.** Medicine is complicated, edge cases are everywhere, and lawsuits ain’t cheap. Getting doctors to toss an agent the keys to complex workflows is a tall order, but that’s exactly why Amigo designed its entire platform around getting that buy-in with verifiable evidence every step of the way. **The Takeaway** Clinical AI has the potential to transform healthcare. Fine-tuned AI agents can help eliminate medical errors, keep patients engaged with their care, and allow providers to start carving out competitive moats through their own clinical differentiation. Doctors aren’t going to arrive at that future by taking a leap of faith. Trust is gained slowly, and can shatter instantly. AI agents will have to earn credibility one workflow at a time, and could lose it all with a single misstep. That said, it’s a future worth striving for, and Amigo’s safety-first approach to building trustworthy AI agents is one of the best roadmaps we’ve seen for how to get there. *Nothing gets the magic across better than Amigo’s live walkthrough. Make sure to check out the agents in action by* [*booking a demo on their website*](https://www.amigo.ai/)*.* --- # AI Doctors Are Like Self-Driving Cars: Lessons from Waymo to Build Trust Source: https://www.amigo.ai/blog/lessons-from-waymo Author: Richard Wang Published: 2025-09-19 Building trust through scoped domains, rigorous simulation, and human-in-the-loop safeguards. Over the past year, Waymo’s momentum has shifted the autonomous vehicle conversation from “if” to “how soon.” The company is rapidly adding cities and integrating with public systems and mainstream partners. Waymo and Via just announced a public-transit integration in Chandler, AZ. And with Lyft, Waymo is building a Nashville robotaxi network. Airport pilots are advancing too, with new permits and testing on California airport grounds. While this feels like it’s happening fast, what we’re actually seeing is the gradual, methodical rollout of a system designed to earn trust by delivering measurable safety. Recent analysis of Waymo's performance shows a stunning 91% reduction in serious crashes compared to human drivers, with 96% fewer injury incidents at intersections (historically the deadliest driving scenarios). This is not incremental harm reduction; it looks like harm prevention, with the potential to eliminate traffic deaths as a leading cause of mortality. The key to Waymo's success is not just their technology. It's their careful approach to earning trust in a high-liability, high-variance domain where every interaction can be uniquely complex, and the cost of failure is measured in lives. And that playbook offers crucial lessons for another life-or-death industry grappling with AI deployment: healthcare. So if you’re trying to build an AI doctor, the core question is the same one Waymo confronted: *How do you guarantee performance?* ### High Stakes, High Variance, High Skepticism Every patient interaction, like every driving scenario, involves complex, dynamic factors that make each situation unique. A slight medication dosage error or missed diagnosis can be as fatal as a miscalculated turn at an intersection. And just like with self-driving, public trust is everything. The safety and performance bar that an AI doctor needs to cross is much higher than the average human benchmark. The question is not whether the technology *can* work, it's whether it *will* work – in ambiguous settings, 100% of the time. This is where most AI companies get it wrong. They pursue the "general-purpose clinical agent," which is the equivalent of launching a self-driving car that claims to work perfectly in every city, weather condition, and traffic scenario from day one. As many healthcare buyers have learned the hard way, these are the agents that demonstrate great potential in demos but break down in production when faced with less-than-perfect real-world conditions. Waymo took the opposite approach, and healthcare AI should follow suit. ### The Waymo Playbook Within the field of autonomous vehicles, an *operational design domain (ODD) *defines the specific conditions under which an AV is designed to operate safely, including environmental factors like weather, geographical limitations, time-of-day restrictions, and roadway characteristics. ODDs are crucial for safety, as they establish the boundaries of the system's capabilities, ensuring it functions reliably within its defined parameters and does not operate outside its design limits. Waymo’s design approach was to master an ODD – say, a neighborhood in Phoenix – under specific conditions, then widen the geofence as competence and confidence grow. They mapped every street, understood every traffic pattern, and simulated countless scenarios within that constrained environment before putting a single car on those roads. Even then, they didn't go fully autonomous immediately. They put safety drivers behind the wheel to monitor performance and intervene when necessary. Only after proving exceptional safety and reliability in that specific domain did they remove the human oversight and expand to new neighborhoods. These five elements are what allowed them to build trust: 1. Domain Specificity 2. Virtual Validation 3. Human Oversight 4. Gradual Autonomy 5. Methodical Expansion Let’s break down how this applies to building clinical AI agents. ![Blog image](https://www.amigo.ai/images/blog/content/lessons-from-waymo-0.svg) ### **1. Domain Specificity: Define a Clear ODD** Amigo takes this approach for healthcare AI. Instead of a general-purpose “AI doctor,” we help provider companies build custom agents for specific clinical neighborhoods. This extends to both different specialties (e.g., women’s health, cardiology, oncology) and different use cases (e.g., triage, post-visit follow-ups, care coordination). Each agent is scoped to the conditions and populations where it can be thoroughly validated. Rather than hoping a general-purpose agent will safely handle the complexities of different medical domains, we work with clinical experts to define an agent's ODD – the precise scenarios where it can reliably operate. ### **2. Virtual Validation: Simulation Before Street Time** Once the agent has been built, Amigo’s partners will have their clinicians and product leads tangibly define what “good” looks like. We then simulate that environment using high-fidelity synthetic patient interactions – hundreds of thousands of them – graded on these bespoke safety and performance metrics (accuracy, appropriateness, empathy, escalation reliability, regulatory adherence, etc.). This exposes rare edge cases in hours that might take years to encounter in the field, all before the agent ever talks to a real patient. The grader in this case is Amigo’s custom-built AI Judge, capable of evaluating agent success in simulations at 100,000x the speed of a human. Crucially, this allows for iterative improvement to be done at scale and significantly shortens time-to-convergence. Naturally, this begs the following question: *How do you trust the Judge?* ### **3. Human Oversight: Teach the Judge to Judge** Waymo's initial rollout included safety drivers who could take control when needed. Similarly, Amigo’s deployment process has clinical oversight built in: first in simulations, then in real-life production. We work with human clinicians to assess the fidelity of the simulated world and the behavior of the agent, but perhaps their most important role is in refining the accuracy of the Judge. This is because the Judge is actually the strongest source of evolutionary pressure in the agent development process. As long as we can correctly define success and train the Judge to adjudicate correctly, we will create a strong safety net to ensure an unfinished agent never leaves the factory. As performance stabilizes and trust in the Judge grows, oversight tapers – not based on general feelings, but by hitting explicit, pre-agreed metric thresholds. ### **4. Gradual Autonomy: Earn the Right to Fly Solo** Just as Waymo eventually removed safety drivers as they proved reliability, we can gradually reduce reliance on human oversight. Once clinicians have gained the confidence that the Judge is calibrated correctly and the agent is performing safely across common and edge cases, Amigo’s system can take over from human reviewers to perform quality assurance at scale. We continue to run the Judge on the same metrics we ran in simulation when the agent is live in production with real patients, continuously monitoring for any signs of degradation. And human clinicians are still able to conduct random quality inspections to ensure there is no drift over time. Letting go of the reins can be the scary part. However, if we’ve brought clinicians along for the whole journey and demonstrated trust and safety at every step, we find that this step becomes easy. ### **5. Methodical Expansion: Onto the Next Neighborhood** Once a clinical agent is deemed to be safe and effective in one “neighborhood,” it becomes ready for scope expansion. As willingness to grow the capabilities of an agent increases, or as regulatory changes occur to create more space for AI in care delivery, we can teach the agent to expand its surface area in a controlled manner that leaves nothing up to chance. This also includes adjacent use cases that may require entirely new protocols and structures. One significant advantage to building agents in this modular manner is that new agents can efficiently inherit validated behaviors and guardrails. The network effect is cumulative: every new scenario adds data and robustness to the whole system, increasing the contextual intelligence of agents across the board. ### Building the Future Responsibly Waymo’s expansion from a few dozen cars in Phoenix to 1M+ monthly rides in cities all over the US didn't happen overnight. It took careful validation, continuous improvement, and a relentless focus on safety. But that measured approach is exactly why they're now the clear leader in self-driving cars, with the safety data to prove it (sorry, Tesla). And this data shows they've achieved something remarkable: the transition from trying to reduce harm to actually preventing it. Their cars have managed to systemically avoid the conditions that lead to crashes altogether. Clinical AI has the same potential. Properly designed AI agents can prevent medical errors from happening and ensure no patient ever falls through the cracks. This could mean anything from providing 24/7 availability for urgent questions to delivering consistent, evidence-based care at scale. But realizing this potential requires abandoning the "move fast and break things" mentality that has worked for AI deployment in other industries. Instead, we need the methodical, safety-first approach that's making Waymo a success. The companies that will win aren't those promising to solve everything immediately, but those building trust through demonstrated competence in specific domains. They're the ones running hundreds of thousands of simulations before real-world deployment, keeping humans in the loop until autonomy is earned, and expanding thoughtfully from proven success. With the capabilities of today’s AI models, the future of healthcare isn't an off-the-shelf general-purpose AI doctor that magically handles everything. It's a network of specialized, highly-trained AI agents – each proven safe and effective in their specific domain – working alongside human clinicians to provide better, more accessible care. To learn more about Amigo’s approach to safety, book a chat with us [**here**](https://cal.com/alikhokhar/amigo-demo). --- # Deep Dive: Functional Memory Source: https://www.amigo.ai/blog/functional-memory Author: Ali Khokhar Published: 2025-09-10 Introducing Amigo's Functional Memory architecture that allows agents to think, learn, and remember like healthcare professionals. *This is the first in a five-part series diving into the Amigo Cognitive Architecture. In upcoming posts we will explore Context Graphs, Dynamic Behaviors, Actions, and the Agent Core.* Traditional AI memory systems are prone to breakdowns at the most inopportune times. They may forget essential context or misinterpret information. In healthcare, where accuracy and timeliness are critical to patient safety, this is unacceptable. Amigo solves this problem with a twofold approach: 1. A *functional memory* system that enables hyper-optimization for different clinical use cases 2. A *layered memory *architecture to prevent information density explosion Let’s take a deep dive into Amigo’s functional memory system and how it solves the challenges unique to healthcare. ### Functional Memory: Remembering What Matters To understand what this means, it’s helpful to describe the dimensional *user model* framework that serves as the blueprint for the functional memory system. The user model is a set of dimensions that determines how to categorize information through a specific clinical lens – including what information to store forever (e.g., family history) vs. capture ephemerally (e.g., casual conversations) and how to preserve and decay each data type appropriately. What a patient’s PCP considers important to remember about them is different from what their orthopedic surgeon considers important, and a single piece of information may be interpreted in different ways. As a result, user model dimensions are designed to be unique to each specialty, service, and even clinic. Through viewing each new piece of information against the holistic understanding of the patient, Amigo’s functional memory system can also recontextualize past data and arrive at novel insights that are missed via purely additive approaches. When a patient is diagnosed with ADHD, Amigo can re-evaluate a six-month-old complaint about concentration difficulties. This re-analysis can surface new patterns that were not apparent during the original interaction. As a result, the agent’s memory becomes more clinically relevant over time. Recontextualization also allows for temporal pattern recognition, allowing the agent to understand how a patient’s health evolves over time and distinguish between temporary pain and chronic condition progression. Another crucial insight to consider is that even the definition of what matters is subject to change, as medical or regulatory stances evolve over time. Amigo’s functional memory system was built to allow for the continuously updating of user model dimensions *without losing fidelity of past memories *– the system can perform a retroactive backfill to reinterpret past information in a new way, ensuring a patient’s history stays relevant as the outside world changes. ### Layered Memory: Managing Information Overload One of the biggest challenges faced in AI memory design is how to handle the sheer volume of data accumulated over time – we need to prevent the loss of key insights while also ensuring that the system is not overwhelmed by large amounts of junk that make it impossible to reason. The human brain does a good job of this by intelligently prioritizing and organizing key insights to remember while systematically decaying information that doesn’t matter. Amigo’s functional memory system relies on a layered architecture to emulate this intelligent approach. Let’s break down each layer. **L0: Raw Transcripts** – Retains every word ever said, providing the foundation for historical recontextualization during live patient interactions and acting as the ground truth for future references. Simple memory systems stop here, but suffer from the *information density explosion* problem described above. **L1: Extracted Memories** – After each conversation, the agent engages in *post-processing* (like dreaming, when the brain organizes relevant information from the events of the day for long-term storage). This process extracts new memories from L0, determines what is novel vs. redundant information, and decides what’s worth keeping long-term. **L2: Episodic User Insights** – Organizes L1 notes into meaningful insights that are structured based on the dimensions of the user model (L3). The process of converting L1 memories into L2 insights forms temporal checkpoints that capture changes in our understanding of the patient over time. This step acts as a layer that prevents information density overload when ingesting insights into the user model. **L3: Global User Model** – This represents the highest-level understanding of the patient as a whole, organized by dimensions that capture the most important information for a specific type of clinician to know about them. We can maintain an up-to-date understanding of the patient across time, providing a clear and complete picture that the agent can draw from during a conversation – this creates 90-95% efficiency gains in active memory retrieval. When more niche information is required but missing, the agent writes its own targeted query to dig deeper and fill in the gap. Rather than broadly searching memory for phrases like “leg pain” to respond to a user query, Amigo’s system contextualizes the question against the user model to ask the question in a smarter way and efficiently extract deeper insights. The ability to flag information gaps also allows the agent to ask follow-up questions in real time. ### Moving from Blueprint to Bedside Healthcare conditions are complex, requiring a synthesis of patient health history, family history, medications, symptoms, age, and multiple additional factors. With the global user model at its core, Amigo agents are equipped with true functional clinical intelligence, acting like skilled clinicians who always have the most relevant information at their fingertips. Let’s recap what we’ve covered: During live sessions, Amigo’s functional memory system keeps the user model (L3) active at all times. This allows agents to reason instantly with the complete patient context, removing the latency-accuracy tradeoff that affects other AI memory systems. It also creates multiple interconnected feedback loops between the global understanding of the patient and the immediate conversation, interpreting every detail based on a provider’s identity and service line. In other words, a cardiology agent interprets chest pain through a cardiovascular risk assessment framework, while a psychiatry agent will consider anxiety manifestation and somatic symptoms. The functional memory system then identifies new information and evaluates it in the context of the patient’s medical history, merging relevant updates into the user model without losing clarity or becoming overwhelmed. Because Amigo agents keep the whole patient’s picture in mind constantly, interpret new information in real time, know what’s truly important, and prevent information overload, they are able to reason like trained medical professionals instead of simply reciting facts. ### Deliver Context-Aware Clinical Intelligence Like a doctor who remembers all the important details of a patient’s past medical history and can put them in context based on the latest information, Amigo’s functional memory system remembers what matters when it matters most. As a result, organizations can deliver personalized care across hundreds of thousands of patients. Have more questions about the Amigo memory system? [Book a call](https://www.amigo.ai/) and our team will be happy to answer them in detail. ![Blog image](https://www.amigo.ai/images/blog/content/functional-memory-0.png) --- # Healthcare AI Infrastructure: Build vs. Buy? Source: https://www.amigo.ai/blog/build-vs-buy Author: Ali Khokhar Published: 2025-08-21 Navigating implementation trade-offs, hidden costs, and architectural complexity in clinical AI deployment. You have a product roadmap filled with aspirational healthcare-specific AI use cases and patients who need care now. Do you build your agentic AI platform in-house, or do you partner with a vendor? What parts should you build and what parts should you outsource? Before you answer the classic build-or-buy question, take a step back and ask yourself a deeper question: ***Do you want to become an AI infrastructure company, or would you rather be the best at delivering AI-powered healthcare?*** The answer will impact your quality of care, competitive positioning, time to market, scalability, and more. Let’s explore why. ### Build or Bust (Too Often, It’s Bust) Many healthcare companies start their AI journey with certainty that building their infrastructure in-house is the smartest move. *Why would I outsource something that is so strategically important?* This concern is perfectly valid. In healthcare, it’s crucial to have the freedom to customize your agent to meet your specific patient populations, use cases, and compliance requirements. However, there is a big difference between *building agents* and *building infrastructure*. A strong technology partner who understands the healthcare space will not sell you agents out-of-the-box. They will build a product specifically designed to give you the tools and levers to execute the level of control you need. The paradox of attempting to build all of your scaffolding in-house is that it actually results in a weaker technical infrastructure with *less* control over agent behavior, not more. Not to mention the desire for perceived control can set your launch back by 18+ months. A second misconception is the belief that owning your own AI infrastructure will give your company a competitive moat. The reality is that when you’re not an AI company at your core, it is virtually impossible to build and maintain a long-term advantage on technology alone. More on this later. Instead, where healthcare organizations have the right to win is developing a patient experience that drives engagement, regulatory compliance that ensures trust, and care protocols that deliver better outcomes. Some teams also fall victim to “not invented here” syndrome, resisting external solutions in favor of homegrown systems. The irony is that this inward focus leaves you reinventing the wheel while your competitors move ahead with safe, responsible, and market-tested AI tools. ### The Hidden Complexity of AI Orchestration It is easy to overlook the layers of complexity needed to create a truly safe agentic AI platform. ![Blog image](https://www.amigo.ai/images/blog/content/build-vs-buy-0.svg) When building trustworthy AI for high-stakes healthcare use cases, the visible outputs *–* from the agent’s responses to the overall product experience *–* rest on a deep, interconnected foundation. Building this foundation requires solving interconnected technical challenges that compound rapidly, making it extremely difficult for organizations without deep AI infrastructure expertise to develop effective solutions. To give you an idea of what this means, here are some of the problems we had to solve along our journey: ### **Core agent development** - **Information Density Management**: How do you handle information density explosion in context windows? - **Session-to-Global Context Integration**: How do you aggregate session-level insights into a coherent global user understanding? - **Full Context Traceability**: How do you trace the full context behind any agent decision? - **Emergent Pattern Integration**: How do you systemically translate observable user patterns into agent improvements? ### **Voice capabilities** - **GPU Procurement and Optimization**: How do you achieve low-latency voice processing at scale? - **Hardware Availability Management**: How do you handle GPU shortages in key geographic regions? - **Regional Compliance Navigation**: How do you handle voice processing when standard providers don't meet compliance requirements? - **Enterprise Infrastructure Scaling**: How do you manage infrastructure costs of hundreds per hour plus $30K+ in licensing? ### **Agent actions** - **Cold Start Optimization**: How do you eliminate cold start latency in action execution? - **Dynamic Compute Scaling**: How do you handle unpredictable variable action workloads? - **Comprehensive Security Scanning**: How do you comprehensively scan action tools for security vulnerabilities? - **Custom Environment Requirements**: How do you build custom compute environments for specialized action requirements? ### **Simulation system** - **Conversation Scope Management**: How do you maintain conversation coherence while covering comprehensive test scenarios? - **Parallel Execution Infrastructure**: How do you run large-scale simulations without hitting provider limits? - **Statistical Coverage Analysis**: How do you measure what percentage of real scenarios your simulations actually cover? - **Production Drift Monitoring**: How do we track gaps between simulated scenarios and actual production behavior? ### **LLM-as-a-judge evaluation system** - **Contextual Evaluation Architecture**: How do you give judges full access to user session history and database context? - **Adaptive Reasoning Systems**: How do you handle complex evaluation cases that require dynamic reasoning depth? - **Multi-Model Orchestration**: How do you orchestrate multiple specialized models for different evaluation scenarios? - **Multi-Provider Load Balancing**: How do you build custom load balancing across regions and providers while maintaining strict compliance? These examples represent just a fraction of the challenges we encountered. Each solution required specialized expertise in machine learning infrastructure, distributed systems, security, compliance, and enterprise scalability. The problems compound—solving one often reveals three more, creating complexity webs that span multiple engineering domains. ### The Hidden Costs of In-House Infrastructure The technical rigor of building an AI solution in-house is just one of the factors at play. Another is cost. And once you add it all up, you’ll find that trying to build from the ground up typically delivers a slower, more expensive path to the same destination. Three areas to consider: ### **Time-to-market realities** Most teams assume they can get an AI clinician up and running in three to six months. A more realistic timeline is 18+ months if you’re tasked with building the architecture from scratch – not to mention the associated costs in engineering talent. Working with a partner can shrink that timeline dramatically, freeing you from infrastructure development and providing you with a fully trained, production-ready agent you can deploy in as little as six-to-eight weeks. ### **Fragmentation costs** When most companies say “build,” they’re actually just buying the foundational components they can’t build from scratch – agent frameworks, memory systems, monitoring tools, evaluation platforms – and stitching them together. But this Frankenstein-style approach creates a monster that’s hard to tame. You could spend months trying to get these disparate tools to communicate, let alone achieve clinical functionality. Along the way, you’ll increase costs, introduce reliability risks, and fall short on safety requirements. ### **Ongoing maintenance burdens** Any homegrown AI orchestration system will require ongoing evaluations and regular model upgrades. And if security and compliance checks lag behind regulatory changes or user expectations, technical debt can build quickly. Even more costly is the need for continuous, resource-intensive testing on edge cases. In these rare scenarios, a false positive or negative could create serious patient harm. ### The Strategic Benefits of Partnership: Speed, Safety, Cost When you factor in the combined hidden costs, building entirely in-house starts to look less attractive. A smarter, more sustainable approach: partnering with an AI solution provider who understands the technology and the unique challenges of ensuring trust and safety at scale. With the right partner in place, your team can stop wrestling with AI infrastructure challenges and focus on building a competitive advantage grounded in clinical differentiation. Your AI partner can then handle the rest. How much faster can a partner move than your in-house team? Here’s what a typical timeline looks like for companies that choose Amigo to deploy AI agents: - **Weeks 1-2:** Agent design and clinical workflow mapping with your team. - **Weeks 3-4:** Agent training and initial testing in our simulation environment. - **Weeks 5-6:** Integration with your systems, plus comprehensive simulations and evaluations. Once fully operational, you can start seeing results and gathering real patient feedback in less than two months. All without spending hundreds of thousands of dollars on expensive in-house AI engineering talent. But not just any partner will do. In healthcare, the cost of making mistakes is extremely high, and AI solutions must be [trustworthy and compliance-ready](https://www.amigo.ai/blog/how-amigo-solves-ais-trust-crisis). Seek solution providers that offer full observability so you can understand the reasoning behind every decision a clinical intelligence agent makes. Also, look for AI companies that rely on [real-world validation](https://www.amigo.ai/blog/beyond-benchmarks-why-healthcare-ai-needs-real-world-validation) rather than standardized performance benchmarks. This approach proves performance in messy scenarios that reflect real patients, such as identifying dangerous drug interactions in a person with multiple chronic conditions. ### Forging a Win-Win Partnership Partnering with a vendor doesn’t mean you have to surrender control of your data or your IP. In fact, the opposite is true. Amigo gives our partners the tools to implement their clinical vision and maintain full control of the technology while leveraging our platform infrastructure. You retain oversight over your AI agent’s behaviors, defining everything from the agent’s clinical protocols and safety boundaries to its conversational tone and evaluation criteria. All of your company’s clinical intelligence – including every insight from your patient interactions – remains entirely under your control. Amigo will never provide agent configurations or training data to other companies. The best partnerships happen when there’s a strong fit on both sides. So when evaluating potential AI partners, ask a few key questions: 1. Will they let your clinical teams lead development? 2. Can they provide transparent and rigorous safety protocols? 3. Can their agents seamlessly integrate with your existing systems? Use the answers to find a partner that will meet your specific requirements. ### Why It Matters The build vs. buy decision in healthcare AI may seem simple, yet it’s anything but. Choosing the wrong path can slow innovation, increase costs, and divert your team’s attention from what really matters: creating better patient outcomes. At Amigo, we understand that engineers don’t know how to build the brain of a clinician. But clinicians shouldn’t need to become AI engineers either. With our infrastructure, your product and clinical teams can build and iterate without requiring engineering expertise. We’ll do the heavy technical lifting so you can focus on getting your safe clinical agents to market faster, create a competitive moat, and ultimately deliver exceptional patient care. Learn more about [trusted AI for healthcare](https://www.amigo.ai/use-cases). --- # Beyond Benchmarks: Why Healthcare AI Needs Real-World Validation Source: https://www.amigo.ai/blog/beyond-benchmarks Author: Ali Khokhar Published: 2025-07-30 Standardized performance benchmarking fails to prepare AI agents for the nuanced situations faced by actual patient populations. Invite a healthcare AI vendor into your conference room and prepare for a barrage of benchmarks. You’ll hear claims of 95% accuracy on medical questions or 90% diagnostic precision on curated datasets. The question is: Can you trust them? The problem isn’t with the benchmarks themselves—which serve as useful starting points—but it’s how they’re built and interpreted. Most vendors optimize for statistical success on average cases, using curated datasets. But healthcare’s most important decisions occur on the edges, with the most complex, vulnerable cases, which is exactly where these benchmarks don’t measure performance. Consider an emergency department that might see hundreds of routine cases for every true crisis. Traditional benchmark testing would create AI tools that excel at treating common cases but miss potentially critical failures in rare but life-threatening situations, such as an anaphylactic medication reaction. Instead of relying on lab-perfect benchmarks rooted in sanitized use cases, healthcare organizations need a better way to test and trust AI solutions. And real-world simulations provide a smarter path forward. Here’s why. ### “Good Enough” Isn’t Enough When other industries fine-tune their products, they use the 80/20 rule. If vendors improve the experience for 80% of their users, it justifies any shortcomings for the other 20%. That might be acceptable for low-stakes use cases like customer service. But in healthcare, these traditional benchmarks are not sufficient. A patient-facing AI agent must account for every potential complex scenario: The confused elderly patient with multiple medications. The anxious patient describing symptoms at 2 a.m. The non-native speaker trying to communicate pain levels. These complex cases require near-perfect performance. Anything less, and your organization could cause harm to its patients. Traditional benchmarks also favor theoretical capability over practical implementation. Yet theory has a well-documented track record of being terrible at predicting real-world performance. Consider the plight of the original AI-powered Epic Sepsis Model. It delivered theoretical accuracy rates between 76% and 83% in development. But in real-world applications, it missed [67% of sepsis cases](https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307), causing Epic to [overhaul the algorithm](https://www.beckershospitalreview.com/healthcare-information-technology/ehrs/epic-overhauls-sepsis-algorithm/) for improved performance. In the interim, organizations relying on the model ran the risk of higher-than-expected patient infection rates. ![Blog image](https://www.amigo.ai/images/blog/content/beyond-benchmarks-0.svg) ### Passing the Test but Failing the Patient Obsessing over lab-perfect benchmarks while evaluating AI solutions is like a doctor or nurse cramming all night to pass their board exam. Their high scores may reflect theoretical knowledge, but it doesn’t make them an excellent clinician. When AI vendors talk about their test scores but ignore outcomes, they miss what truly matters in patient care: trusted, validated clinical results. Vendors who laser-focus on curated test scenarios also inevitably end up trying to game the system, tuning models to improve scores without addressing critical edge cases. Traditional benchmarks also give vendors and healthcare organizations a false sense of security. Few solution providers track performance drift, ensuring, for example, that a medical AI agent still provides correct drug dosages after an update. And in healthcare, even well-designed benchmarks become outdated quickly as medical knowledge evolves. When the evaluation criteria lags behind current clinical practice guidelines, AI agents begin drifting further from relevance. ### Understanding the High Stakes The impact of a healthcare AI solution validated by lab-perfect benchmarks ripples throughout an entire healthcare organization. Vulnerable patients are the first to suffer. Consider a teenager who downplays the seriousness of his depression symptoms. An AI system tested with traditional benchmarks could miss subtle nuances within his answer, leading to inadequate mental health support and a longer recovery. For providers, relying on AI models evaluated against generalized statistical targets can lead to inaccurate diagnoses. A model trained to address urgent care concerns, for example, won’t work for a cardiology company. Just as physicians continue learning after medical school by performing residencies in their specialty, AI models need to be tested in environments that reflect the realities and risks of the care settings it’s built to support. At a system-wide level, healthcare needs trusted evaluation standards that predict clinical success at the highest level of accuracy. Organizations that master evidence-based validation will transform patient care, while those mired in traditional benchmark thinking will struggle to keep pace. ### The Critical Importance of Real-World Simulations Simulation testing bridges the enormous gap between lab performance and real-world effectiveness by creating sophisticated environments that reveal a healthcare AI agent’s true operational readiness. At Amigo, we understand both AI and the unique challenges of delivering quality patient care at scale. Our focus on simulations helps healthcare organizations find what lab-perfect benchmarks miss. Our evaluations platform builds a parallel universe where we stress-test your agent against the same challenges it will encounter in the wild, but in a controlled environment where every action can be measured and analyzed. Amigo trains its agents in simulated environments that reflect the real-world scenarios and demographics unique to each customer’s patient population. We also tailor our evaluation rubrics to specific clinical roles, defining success differently for a nurse triaging a patient versus a specialist making a diagnostic call. The overarching goal: to push AI to the edge, break it, fix it, and then verify it works under pressure, so patients and providers can trust it. We deliberately oversample statistically rare scenarios—like a patient with unusual drug interactions—because we know their importance far outweighs their frequency. This importance-weighted testing helps keep our agents consistent at scale. Rather than relying on human reviewers whose standards might vary with fatigue or mood, our AI judges evaluate every interaction against precise criteria. These judges receive up to 50 times more computational resources than the agents they evaluate, allowing them to determine whether each interaction delivers the value your organization promises. ### How We Evaluate Clinical Intelligence Amigo shifts the AI healthcare conversation from statistical promises to evidence-based confidence. In well-defined areas like prescription verification, agents may achieve 99.9% accuracy quickly, giving you the assurance to deploy them. Other more nuanced tasks, like mental health support or crisis detection, may take longer to reach this threshold. Our evaluations system provides an honest assessment of which scenarios the agent is handling well vs. those it might miss, so that targeted development and testing can be done before real-world deployment. Once we deploy an agent, we continue stress-testing it so we can detect degradation early, minimize drift, and intervene long before it impacts patient care. Our deployment architecture includes real-time monitoring, allowing us to pinpoint improvements for specific issues—such as an error in dosage suggestion—without requiring a full-scale update. ### Building Trust That Matters While trust is built through real-world testing, it’s maintained by staying in sync with evolving healthcare regulations and changing patient behaviors. From a compliance standpoint, Amigo tracks regulatory mentions, flags interactions that might require updated requirements, then calculates the risks of operating with outdated understanding. If updates are needed, such as revisions reflecting anticipated changes to the [HIPAA Privacy Rule](https://www.hhs.gov/hipaa/for-professionals/regulatory-initiatives/index.html), we can make updates based on those requirements instead of having to retrain the underlying models. Amigo also tracks patient expectations, monitoring survey completion rates and emerging complaint patterns, and adjusts the agent accordingly. Continuous feedback loops between real-world events and your simulated environment allow Amigo to analyze changing patient demographics and suggest new personas to address potential gaps. This advanced capability enables organizations to determine whether a new simulated persona, such as a 35-year-old gig worker juggling multiple chronic conditions with inconsistent insurance, could help service an evolving patient population. And when it comes to patient safety, Amigo uses comprehensive regression detection to catch subtle degradations before they turn into serious problems. Instead of wondering whether medical AI still provides correct drug dosages after an update, our clinical intelligence platform validates it while also suggesting potential changes to patient communication strategies. ### Embrace an Evidence-Based Approach to Healthcare AI Operating on faith—backed by lab-perfect benchmarks—won’t help your organization implement AI safely and securely at scale. Instead, operate on evidence. Solution providers that stress-test their agents in real-world simulations can help you deploy AI confidently, scale it wisely and embrace continuous improvement to benefit your providers and patients. To learn more about Amigo's approach to evidence-based testing, [**book a time with me here**](https://cal.com/alikhokhar/30min). --- # Amigo Partners with Eucalyptus to Scale AI-Powered Healthcare Delivery Across Global Telehealth Network Source: https://www.amigo.ai/blog/amigo-partners-with-eucalyptus Author: Richard Wang Published: 2025-07-28 Strategic partnership aims to deploy AI health assistants across Australia's largest telehealth platform serving 200,000+ patients in four countries. *Click *[*here*](https://www.prnewswire.com/news-releases/amigo-partners-with-eucalyptus-to-scale-ai-powered-healthcare-delivery-across-global-telehealth-network-302514630.html)* to read the following press release on PR Newswire.* --- NEW YORK, July 28, 2025 /PRNewswire/ -- [Amigo](https://c212.net/c/link/?t=0&l=en&o=4475152-1&h=3880486033&u=https%3A%2F%2Fwww.amigo.ai%2F&a=Amigo), the AI infrastructure company building trustworthy agents for high-stakes environments, today announced a strategic partnership with leading global telehealth provider [Eucalyptus](https://c212.net/c/link/?t=0&l=en&o=4475152-1&h=507664916&u=https%3A%2F%2Fwww.eucalyptus.health%2F&a=Eucalyptus). The multi-year collaboration will integrate AI-powered health agents into Eucalyptus's portfolio of specialized healthcare brands, beginning with the successful deployment of an AI health assistant for the company's Juniper weight management clinic. This partnership represents a significant step forward in addressing the critical challenge of staff shortages within the healthcare industry, where well-timed engagement often determines patient outcomes. By providing continuous, intelligent support, the AI health assistant helps maintain patient motivation and adherence to treatment plans while capturing valuable insights that inform clinical decision-making. "We are committed to pioneering the future of digital healthcare," said Benny Kleist, Co-Founder and Chief Strategy Officer of Eucalyptus. "Working with Amigo represents a significant step forward in our mission to make quality healthcare more accessible, allowing our clinical teams to ensure every patient receives the always-on, personalized support they need to achieve their health goals." The initial success of June, achieving a \<2% human escalation rate with 100% performance on clinical safety metrics in live deployment, establishes a new benchmark for AI care delivery and validates the potential for scaled deployment across multiple therapeutic areas. Following this deployment, Amigo and Eucalyptus plan to expand AI-powered healthcare support across Eucalyptus's broader network of specialized virtual clinics, each tailored to address specific therapeutic areas and patient populations. To learn more about the deployment of June, Juniper's AI health assistant, please visit [**this link**](https://c212.net/c/link/?t=0&l=en&o=4475152-1&h=3413872703&u=https%3A%2F%2Fwww.amigo.ai%2Fcustomers%2Feucalyptus&a=this%20link). Media Contact: Richard Wang richard@amigo.ai **About Amigo** Amigo builds AI infrastructure designed to solve the trust crisis preventing widespread AI adoption in high-stakes environments. The company has developed proprietary agent architecture that enables enterprise organizations to safely create, train, and deploy AI agents built around three core principles: controllability, continuous alignment, and real-time observability. Amigo's mission is to enable the transition from AI as experimental technology to trusted infrastructure for the economy's most critical functions. For more information, visit [Amigo's](https://c212.net/c/link/?t=0&l=en&o=4475152-1&h=967732561&u=https%3A%2F%2Fwww.amigo.ai%2F&a=Amigo's) website or follow their page on [LinkedIn](https://c212.net/c/link/?t=0&l=en&o=4475152-1&h=3083933388&u=https%3A%2F%2Fwww.linkedin.com%2Fcompany%2Fyouramigo-ai%2F&a=LinkedIn). **About Eucalyptus** Eucalyptus builds and operates a house of digital healthcare companies. Founded in 2019 by Tim Doyle, Benny Kleist, Alexey Mitko and Charlie Gearside, it is now Australia's largest digital health provider and one of Australia's fastest growing companies. Eucalyptus has facilitated over one million consultations across Australia, the UK, and Germany, and is the only telehealth company in Australia certified by the Australian Council on Healthcare Standards (ACHS) against the EQuIP6 standards for safety and quality of clinical services. For more information, visit [Eucalyptus's](https://c212.net/c/link/?t=0&l=en&o=4475152-1&h=1906164371&u=https%3A%2F%2Fwww.eucalyptus.health%2F&a=Eucalyptus's) website or follow their page on [LinkedIn](https://c212.net/c/link/?t=0&l=en&o=4475152-1&h=4118004577&u=https%3A%2F%2Fwww.linkedin.com%2Fcompany%2Feucalyptusvc%2F&a=LinkedIn). SOURCE Amigo Inc --- # Amigo Deep Dive: Healthcare AI Guy Source: https://www.amigo.ai/blog/healthcare-ai-guy Author: Healthcare AI Guy Published: 2025-07-17 Healthcare AI Guy interviewed Ali Khokhar, Amigo's CEO, about how they're using simulation to build trust, turning agents into action-takers, and redefining what build vs. buy means in healthcare AI. *The following is a direct transcription of the interview. To read this post on Healthcare AI Guy's website, please visit *[*this link*](https://www.healthcareaiguy.com/p/company-deep-dive-amigo)*.* --- ## Company Deep Dive: Amigo *Perspectives from the people building the future of health AI…* ![Blog image](https://www.amigo.ai/images/blog/content/healthcare-ai-guy-0.png) We sat down with [Ali Khokhar](https://www.linkedin.com/in/khokharali/?utm_source=www.healthcareaiguy.com&utm_medium=referral&utm_campaign=company-deep-dive-amigo), Co-Founder and CEO of Amigo, a platform for building safe, reliable AI agents in healthcare. Rather than building a one-size-fits-all "AI doctor," Amigo is focused on providing the infrastructure and tooling to help healthcare organizations create their own highly customized agents, from virtual clinicians to care coordinators, with safety, transparency, and control built in. Ali shared how the team is approaching trust, why they see growing demand for patient-facing AI, and how the company is planning for a future where AI agents become trusted, tested pillars of the care delivery ecosystem. **Let’s start from the top. What is Amigo and what problem are you solving?** Amigo is a platform for deploying AI agents in high-risk industries, with healthcare as our primary focus. Everyone is talking about agents, but in healthcare, you can’t afford failures. Our goal is to make these agents trustworthy enough to operate in environments where the cost of error is very high. That means providing organizations with the infrastructure to build, train, test, and monitor agents they can trust, starting with clinical use cases. We don’t provide an off-the-shelf “AI doctor." Instead, we help companies build their own clinical agents, tailored to their needs. For example, an AI clinician built for an urgent care provider looks very different from one built for a women’s health clinic. Amigo is the infrastructure that makes that customization and control possible. **How do you define trust when it comes to AI agents in healthcare?** To us, trust means confidence that an agent will behave reliably in the way you want it to. That confidence depends on three things: control, alignment, and observability. *Control* means being able to shape and constrain the behavior of the agent as the clinical expert. *Alignment* is about making sure the agent acts according to evolving expectations and regulatory needs. And *observability* means you can monitor and understand the agent's decisions in real time. When these three are in place, we believe you're in a position to trust an agent. **What does that look like in practice when someone builds an agent using Amigo?** We guide each of our partners through a structured process. We start by working with a clinical expert and a product lead to define the agent's "operable neighborhood"—the set of scenarios where it can safely operate. Then we simulate that environment and stress-test the agent using synthetic patient interactions, millions of times. This allows us to ensure reliability and performance in an environment that most accurately reflects that specific partner’s real patient population. After deployment, we keep that simulation loop running. When an agent performs poorly or encounters something new, we generate more simulated scenarios and retrain it. It’s like Waymo: you don’t deploy to every city at once. You start in neighborhoods where you have confidence, then expand into new ones—safely. **How do you ensure clinical teams are comfortable adopting AI agents?** It all starts with involvement. The clinicians are co-creators. When we work with a healthcare organization, we pair one of our agent engineers with a clinician and a product lead from their team. That trio defines how the agent should behave, where it should operate, and what success looks like. By keeping the clinical voice in the loop from day one, we build trust and accountability. When it comes to successful adoption, cultural readiness is just as important as technical performance. And when clinicians help train and test the system, they’re much more confident putting it in front of their patients. And by testing the agent via millions of simulated conversations, clinicians gain the confidence that it will perform really well before it enters the real world. **What are some of the most common use cases you’re seeing?** A lot of our partners are building agents to help overworked care teams. Think: 24/7 support for follow-up questions, intake and triage, lab result debriefs, medication guidance, or just general care navigation. We also support fully conversational agents that can perform actions, not just offer advice. Through our Amigo Actions tooling, agents can order labs, write to the EMR, or message other members of the care team. This helps our partners build complex clinical workflows that save their clinicians a lot of time. ![Blog image](https://www.amigo.ai/images/blog/content/healthcare-ai-guy-1.png) **What makes Amigo different from traditional "build or buy" approaches?** We offer something in between. You’re not buying a rigid, off-the-shelf agent. But you’re also not hiring a team of ML engineers to build everything from scratch. Instead, you get the tooling, memory system, reasoning engine, and simulation framework to build your own agent, with full control, fast iteration, and measurable performance. We think the right way is to build *at the right layer*. You focus on determining the agent’s medical reasoning and clinical behavior; we handle the orchestration, infrastructure, and safety tooling. Healthcare organizations are experts at providing high-quality care, not at building complex AI architecture. **What’s under the hood? Are you training your own models?** We don’t train our own models. Frontier AI labs have already invested billions of dollars into training their foundation models and they already contain all the specialized knowledge and medical reasoning capabilities needed. There’s a whole body of academic research that shows this. The real gaps that need to be bridged are *control* and *trust*, and our approach focuses on correctly activating this knowledge and reasoning. To do this, we built a cognitive architecture that sits on top of the foundation models, with custom systems for memory, reasoning, and behavior adaptation. We orchestrate multiple models in real time depending on the task. One model might be used for clinical reasoning, another for empathetic response generation, and another for knowledge retrieval. This lets us route to the best tool for the job. The result is better performance, more control, and lower latency. **What metrics matter most when evaluating AI agents for healthcare?** There’s no one-size-fits-all benchmark. We work with each of our partners to define a "success scorecard" based on their clinical workflows. That includes both safety metrics like accuracy, clarity, and handoff reliability, and experience metrics like tone, empathy, and response time. When we iterate on an agent, we don’t stop until it matches or exceeds human performance in that setting. Our internal simulator agents are adversarially testing and actively try to surface edge cases—for this reason, one of our recent partners saw their agent perform even better in the real world than in simulations. **What’s your long-term vision?** We believe we're moving from a human-based economy to an agent-based one. But to get there, we need infrastructure that verifies performance, ensures safety, and builds trust. In healthcare, that means AI clinicians need to be "credentialed" in the same way human physicians are. That’s what we’re building toward: the verification layer for the agent economy. **Any final thoughts?** Much of the industry is still focused on back-office automation and scribes. We think the bigger opportunity is in frontline patient-facing care, and we’re already seeing it work. The biggest blocker isn’t the tech. It’s the assumption that there’s no safe way to do this. There is. And the opportunity for impact is tremendous. ![Blog image](https://www.amigo.ai/images/blog/content/healthcare-ai-guy-2.jpg) ## Healthcare AI Guy Summary *What stood out, what’s tricky, and why it matters…* Amigo is building the infrastructure for safe, scalable AI agents in healthcare. Instead of launching a one-size-fits-all “AI doctor,” they’re helping health orgs build their own custom agents, from virtual clinicians to care coordinators, trained and tested to perform safely in high-stakes, patient-facing roles. Each agent gets simulated before it goes live. Clinical teams define what “good” looks like, Amigo builds a synthetic environment to match it, and the agent gets trained and tested there until it hits the mark. That same loop keeps running post-deployment. When new edge cases show up, the agent retrains and improves automatically. It’s a different take on AI enablement. While most of the market is still focused on scribes and billing tools, Amigo is betting that the real value lies in front-office care (triage, navigation, follow-ups) and that organizations want control, not pre-built bots. We’re excited about this bold vision, and with a $6.5M seed round co-led by General Catalyst and GSV Ventures, Amigo’s already shifting that future into gear. **What stood out** - **Agents get real-world reps before going live:** Before an agent talks to a single patient, it’s tested across thousands of simulated conversations modeled after real-world cases. These aren’t generic benchmarks. They’re created in collaboration with each customer’s clinicians. - **It’s not advice-only:** With Amigo Actions, agents can now do more than talk. They can order labs, write to the EMR, and route issues to care teams when permissions allow. That’s a meaningful step beyond most “chatbot” systems. - **Retraining happens on the fly:** If an agent performs poorly or gets pushed out of scope, Amigo automatically generates new training scenarios and loops them into the simulation. The system improves itself, without waiting for manual red-teaming. - **Build meets buy:** Health orgs don’t just install an agent and hope it works. They define how it behaves, what its boundaries are, and when to escalate. Amigo provides the infrastructure and orchestration to make that kind of control feasible. - **They’re serious about verification:** The long-term vision isn’t just deploying agents—it’s credentialing them. Think: infrastructure that verifies, tracks, and validates agent performance like a digital version of board certification. **What’s tricky** - **Frontline care raises the stakes:** Back-office AI can be error-tolerant. Patient-facing agents can't. The margin for error is smaller, and the bar for safety, transparency, and oversight is much higher. - **Adoption still hinges on trust:** Amigo embeds clinicians in the training process, but even then, it takes time to change mindsets. Clinicians need to feel like co-owners, not just testers. - **Customization adds pressure:** Because every agent is different, the success of each deployment depends on how well a customer executes. That’s a strength, but also a risk, especially if orgs underestimate the lift. **Final thoughts** Amigo is going after one of the hardest but highest-leverage problems in healthcare AI: building agents that actually deliver care. Their infrastructure-first approach lets customers move faster without cutting corners. If they can keep proving that agents can be safe, transparent, and controllable, they won’t just help the market adopt AI; they might help redefine what trustworthy AI in healthcare looks like. --- # Preventing Clinician Burnout through AI-Powered Care Source: https://www.amigo.ai/blog/preventing-clinician-burnout-through-ai-powered-care Author: Ali Khokhar Published: 2025-06-03 How intelligent patient engagement transforms the care experience from first contact to clinical outcomes. ### When Demand Exceeds Clinical Capacity Many of the healthcare organizations we’ve spoken to describe clinician retention and staffing shortages as their #1 challenge. According to the [latest survey](https://www.ama-assn.org/practice-management/physician-health/national-physician-burnout-survey) by the American Medical Association, over 45% of physicians report symptoms of burnout. Clinical teams are overwhelmed and on top of this, patients can't get the care they need when they need it. Picture the moments when patients need care most: a parent wondering if their child's cough needs medical attention, an elderly patient unsure about mixing their medications, or a diabetic individual noticing unusual symptoms and needing guidance. These patients reach out to their care provider only to encounter lengthy hold times or chatbots incapable of providing advice they can trust. Critical moments like these demand nuanced clinical assessment, empathy, and sound medical reasoning—yet the point of contact leaves patients frustrated and potentially at risk while creating administrative burdens that overwhelm clinical staff. Now for the first time in healthcare history, we can build trusted patient-facing AI agents to manage not just simple intake but clinical care, interpreting medical knowledge through the lens of the patient’s medical history and demonstrating genuine empathy while doing so. These medically-specialized agents can independently resolve the majority of clinical inquiries, intelligently escalating only cases that require human intervention. This breakthrough offers healthcare organizations a way to provide comprehensive patient care at the point of entry, dramatically reducing clinical burden while ensuring patients receive personalized on-demand care. ![Blog image](https://www.amigo.ai/images/blog/content/preventing-clinician-burnout-through-ai-powered-care-0.webp) ### The Patient Journey: From Initial Contact to Resolution The diagram above illustrates an example of how Amigo agents can transform the patient experience from first touchpoint through care delivery. When a patient initiates contact, the custom agent immediately accesses both its own memory of the patient and the provider’s EHR systems to understand their complete medical context and history. The agent can then categorize the patient's inquiry into two distinct pathways. Medical queries trigger a sophisticated clinical assessment process where the agent can provide personalized care based on protocols dictated by experts. For cases that require a human clinician, the system seamlessly escalates to clinical staff with a complete summary of the patient's situation and relevant context. Non-clinical queries follow a separate pathway, where the agent has autonomy to take actions directly, such as updating insurance information, scheduling appointments based on availability and clinical priority, and resolving billing or logistical questions without involving clinical staff. Throughout both pathways, the agent continuously updates patient records and maintains synchronization with existing EHR systems, ensuring that all interactions become part of the patient's permanent medical record. This integration means that whether a query is resolved by the agent or a human clinician, the complete context is preserved for future reference and continuity of care. Below are some examples of how our agents are currently being used in patient-facing clinical deployments: - **24/7 Care and Triage**: Patients receive immediate assessment and guidance as soon as they report symptoms. The agent evaluates symptoms, considers medical history, provides evidence-based care recommendations, and escalates when appropriate. This alleviates burnout for nurses and clinical admin staff while ensuring patients receive timely, appropriate care. - **Pre-Visit Intake and Preventive Care**: Agents streamline the patient experience by collecting comprehensive health information before appointments, identifying potential complications or contraindications, and ensuring patients arrive prepared. They also proactively engage patients for routine screenings, vaccination reminders, and wellness check-ins based on their medical history and risk factors. - **Patient Monitoring and Medication Management**: Agents provide continuous support for patients requiring ongoing monitoring, from post-operative recovery to chronic disease management and cancer care. They track reported outcomes, send medication reminders, provide guidance on drug interactions and side effects, and alert care teams when symptoms or adherence issues indicate the need for clinical intervention. - **Care Navigation and Coordination**: Agents guide patients through complex healthcare journeys, from understanding referral requirements and insurance authorizations to coordinating care across multiple specialties. They help patients navigate treatment options, connect with appropriate resources, and ensure continuity throughout their care experience. ![Blog image](https://www.amigo.ai/images/blog/content/preventing-clinician-burnout-through-ai-powered-care-1.svg) ### Impact on Healthcare Operations Organizations deploying Amigo agents as their patients’ digital front door see substantial improvements across key performance indicators: 1. **Operational Efficiency**: Patient inquiry volumes requiring human intervention drop by 80-90% as agents resolve routine medical and administrative questions directly, allowing clinical staff to focus on high-complexity cases requiring human oversight. 2. **Patient Access**: Response times for common inquiries shift from hours to seconds, with 24/7 availability ensuring patients receive immediate clinical guidance when they need it most, regardless of provider availability. 3. **Clinical Quality**: Enhanced triage accuracy and improved patient education result in more productive patient encounters, better care coordination, and measurable improvements in treatment adherence and health outcomes. 4. **Workforce Retention**: Clinical teams experience reduced burnout as they can free up their time to focus on meaningful patient care rather than an endless stream of routine inquiries, leading to improved job satisfaction and lower turnover rates. ### Partner with Amigo: Building Your Digital Front Door The transformation of healthcare through intelligent, scalable patient engagement is not a distant possibility; it's happening now. At Amigo, we are the first to develop a platform for deploying statistically trusted patient-facing agents in clinical environments, and we are actively partnering with forward-thinking healthcare organizations to create more accessible, efficient, and effective care experiences. To explore how Amigo's AI clinicians can address your organization's patient engagement challenges and reduce clinical burden, please [**contact our team**](https://cal.com/alikhokhar/30min). --- # Evaluations as the Path to Trust Source: https://www.amigo.ai/blog/evaluations-as-the-path-to-trust Author: Ali Khokhar Published: 2025-05-28 How we designed an evaluations system to achieve 99.9% safety scores in high-stakes healthcare environments. In our [**previous post**](https://www.amigo.ai/blog/how-amigo-solves-ais-trust-crisis), we explored why trust is the critical barrier for widespread AI adoption. Enterprises seek confidence that AI systems reliably reflect their goals, values, and priorities. The first step to achieving this confidence is to tangibly define exactly what successful behavior looks like. This is a significant challenge for most organizations, particularly when building expert agents in domains like healthcare, where countless unique patient scenarios and interactions can occur. Clearly articulating measurable success criteria in these complex, unpredictable contexts is essential to reliably evaluate AI performance. Furthermore, as organizational priorities shift and market conditions evolve, these definitions of success must also adapt, demanding an evaluation framework flexible enough to continuously verify alignment and ensure lasting trust. ### Introducing the Amigo Arena To address this challenge, we have developed a rigorous training gauntlet we call the Arena—a structured environment that continuously applies pressure to guide agents toward safe and desirable behavior. The Arena encompasses four key components: 1. Multidimensional Metrics 2. Personas and Scenarios 3. Programmatic Simulations 4. Continuous Improvement Together, these components form an iterative loop designed to systematically build and maintain trust in AI systems. ![Blog image](https://www.amigo.ai/images/blog/content/evaluations-as-the-path-to-trust-0.svg) ### Multidimensional Metrics: Defining Tangible Success The first step involves clearly defining measurable metrics. These metrics translate qualitative expert judgments into quantifiable, objective success criteria. For instance, rather than instructing an AI doctor to "demonstrate good bedside manner," we define specific behaviors—within areas like accuracy in medical diagnoses or clarity in patient communication—that can be consistently measured across millions of interactions. Critically, these metrics are defined by our partners’ clinicians—these medical experts understand patient needs, medical subtleties, and the ethical considerations necessary to define meaningful and relevant evaluation criteria. Conventional measurement systems test one simple metric at a time, often optimizing for academically-defined AI performance benchmarks. In reality, clinical scenarios contain many interrelated factors: medical accuracy, empathy, guideline adherence, risk assessment, and more. For this reason, we built our metrics system to measure holistic outcomes that balance all these critical dimensions, ensuring agents perform effectively in the reality of healthcare interactions. *For example, a sample set of safety & compliance metrics:* ![Blog image](https://www.amigo.ai/images/blog/content/evaluations-as-the-path-to-trust-1.png) ### Personas and Scenarios: Realistic Testing Grounds Metrics alone aren't sufficient without a rigorous environment for testing. This is where simulations come in. We build the guardrails for comprehensive, realistic simulations that mimic the complexity of real-world interactions. Each simulation incorporates: 1. Authentic personas: Detailed representations of the people who will interact with the agent 2. Precisely crafted scenarios: Designed to explore challenging situations and edge cases Each persona is paired with multiple scenarios, creating a comprehensive persona/scenario matrix. Each pairing in this matrix will be re-run across conversational variations to stress test robustness under different conditions, ensuring agents aren’t able to pass tests by chance. The result is a much more comprehensive assessment of agent capabilities, providing clear areas for targeted improvements. *For example, a high-level summary of a simulation set:* ### Programmatic Simulations: Objective and Scalable With clearly defined metrics to measure success and specific personas and scenarios to conduct simulations, we conduct adversarial testing at scale through advanced, reasoning-powered evaluation models. Agents are rigorously challenged, exposing vulnerabilities and enabling iterative improvements. These evaluations measure performance across thousands of simulated interactions to produce a statistically significant confidence score. Patterns can then be visualized via capability heat maps and performance reports. Our evaluators transparently display their reasoning, allowing domain experts and safety teams to audit the logic behind each assessment. This transparency helps identify and correct misalignments quickly, fostering trust and ensuring evaluations remain firmly grounded in professional standards. In conjunction with human testing to provide oversight, programmatic evaluations provide objective insights on safety and performance at full deployment scale. We equip our simulation and evaluation models with 10-50× more AI reasoning tokens than the main agent to enable deeper analysis. These high-compute evaluators avoid the challenge of low-quality data that is common in traditional evaluation systems by running on the organization’s own data that captures their expertise, priorities, and edge cases. The result is a virtuous cycle where evaluations become increasingly precise and relevant over time, even in specialized domains where high-quality training data has historically been limited. *For example, results from a sample programmatic test run:* ### Continuous Improvement: Iterative and Adaptive Good metrics need to adapt as new scenarios and priorities emerge, so they can maintain relevance and precision over time. To this end, the final component of our system is a structured cycle of ongoing measurement, analysis, and refinement. At regular intervals, the complete test set is re-run, ensuring consistent and current evaluation of AI agent performance. Results from these simulations are methodically analyzed against established performance baselines and strategic targets to pinpoint areas requiring attention. After targeted enhancements are made, subsequent evaluations verify whether these enhancements have effectively improved agent performance. Trend analysis reports, improvement tracking dashboards, and business impact assessments are provided to give continuous visibility into progress. This disciplined, data-driven cycle ensures that the agent consistently evolves to meet and exceed organizational objectives over time. And when performance improvements plateau, our reinforcement learning pipeline takes over to push the agent past human ceilings. ### Building Lasting Trust Trust in AI is built gradually, strengthened each time an agent demonstrates alignment with organizational values. The Amigo Arena is designed with this goal in mind: it systematically verifies and improves AI performance in a realistic, measurable, and transparent manner. By clearly defining success through tangible metrics, rigorously testing agents against authentic personas and scenarios, running simulations at scale, and continuously iterating based on data-driven insights, organizations can confidently rely on their agents to not only to meet today's standards but to adapt and grow as expectations evolve. If you’re interested in learning more about Amigo’s evaluations framework, feel free to check out our [**Documentation**](https://docs.amigo.ai/) or [**schedule a call**](https://cal.com/alikhokhar/quick-chat) today. --- # How Amigo Solves AI's Trust Crisis Source: https://www.amigo.ai/blog/how-amigo-solves-ais-trust-crisis Author: Ali Khokhar Published: 2025-05-06 The trust problem prevents AI adoption in critical systems. Amigo builds responsible AI that performs reliably in high-stakes environments. ### The Trust Problem We stand at the threshold of an unprecedented technological shift. Soon, AI agents will become essential parts of our economy—performing complex knowledge work, handling transactions, and managing operations. This integration of AI systems into our economic fabric will make expertise more widely accessible, overcoming constraints on high-skill services previously limited by human capacity. This proliferation of intelligence will transform quality of life, expanding specialized expertise from the few to the many. Despite the enormous potential, widespread adoption of AI faces one critical barrier: *trust*. Organizations hesitate to implement systems they cannot confidently train, control, and audit, and solving the trust problem represents the single most important challenge for meaningfully integrating AI into the economy. We define ***trust*** as confidence that an AI system will reliably and consistently act in alignment with an organization's goals, values, and priorities. This trust is built upon three foundational pillars: 1. **Controllability**: The ability for humans to train, adjust, and intervene in agent behavior to ensure actions remain within acceptable parameters. 2. **Continuous Alignment**: The capability of agents to adapt to changing organizational priorities and maintain goal coherence across different contexts and timeframes. 3. **Real-time Observability**: The transparency of agent operations, allowing organizations to monitor, understand, and verify agent behavior and decision-making processes. Our mission is clear: **build safe, reliable AI agents that organizations can genuinely depend on**. We’ve developed our own agent architecture and built a platform to allow enterprise organizations to safely create, train, and deploy agents into the economy. ### Who We’re Building For The true frontier lies in high-stakes intelligence: AI agents that can operate reliably in environments like healthcare, legal, and finance, where precision and dependability are absolutely essential. High-stakes AI that delivers verified reliability creates outsized value while use cases with a lower performance threshold become increasingly commoditized. Healthcare stands at the forefront of our focus today. The challenge is clear: doctors, nurses, and clinical staff everywhere are overwhelmed by the volume of patient interactions each day. Healthcare professionals are stretched thin trying to provide attentive, personalized care while simultaneously managing complex diagnoses and treatment plans. This cognitive overload leads to clinician burnout and creates barriers for patients seeking timely, affordable, and high-quality care. Consider the complexity of clinical decision support, where AI must navigate thousands of potential diagnoses, remember patient history, and decide on treatment protocols with absolute precision; these scenarios demand a level of reliability that conventional AI systems cannot provide. By building trustworthy AI agents that clinicians can confidently rely on, we aim to alleviate their burden while making quality healthcare more accessible for patients everywhere. ![Blog image](https://www.amigo.ai/images/blog/content/how-amigo-solves-ais-trust-crisis-0.svg) ### The Amigo Architecture: Overcoming LLM Limitations To appreciate how Amigo's architecture solves the trust problem, it's essential to understand the fundamental constraints of large language models (LLMs). At the core of current LLMs is a severe information processing limitation we call the *token bottleneck*. Imagine a human writer who thinks carefully about what to write, inputs one keystroke, suffers sudden amnesia, rereads their document and rethinks carefully, then inputs the next keystroke, and repeats this cycle indefinitely. Dropped reasoning threads, hallucinated details, and occasional nonsensical outputs become inevitable, limiting how much you can trust these systems. Amigo's architecture directly overcomes this limitation through several core innovations: 1. **Context Graphs as Topological Fields**: An expert navigating a problem space is like a rock climber navigating a route up a mountain. By providing the AI models with the right context at the right time, Amigo’s context graph structure unlocks high performance and control by mapping a field of virtual ‘footholds’ that enable the agent to find the path of least resistance. This structure also allows for observable reasoning that humans can inspect and understand. 2. **Functional Memory System**: What your doctor needs to remember about you is different from what your lawyer needs to remember about you. Amigo’s layered memory architecture identifies *what* information deserves perfect preservation, *how* to maintain contextual relationships over time, and *when* to recontextualize information based on new understanding. This allows the agent to reason over the right information at the right density without overwhelming the token bottleneck. 3. **Contextual Knowledge Priming**: A physician who is interpreting a patient's symptoms without relevant context may make an incorrect diagnosis even if they possess all the necessary medical knowledge. What matters is not just having knowledge but activating it in the right context. Amigo’s knowledge system improves through contextual priming rather than perpetually increasing information density, enabling the agent to overcome token constraints and dramatically increasing accuracy. 4. **Targeted Adversarial Testing**: We use multi-turn Judge and Tester agents that punishingly challenge and evaluate performance against custom scenarios and metrics, creating a feedback loop that maintains alignment with expert human judgment. Each edge case or unexpected response becomes an opportunity to immediately correct and improve the agent. Through this model, organizations can indefinitely stress-test and improve their agents until they reach an uncompromising performance threshold. Iterative training achieves *trust as a constant*, not as a one-time exercise. To learn more about Amigo’s architecture, please visit our [**Documentation**](https://docs.amigo.ai/). ![Blog image](https://www.amigo.ai/images/blog/content/how-amigo-solves-ais-trust-crisis-1.svg) ### Speed as Competitive Advantage We designed the system from the ground up with a focus on three decisive time-based advantages: 1. **Time to Trust:** Our approach reduces verification timelines from months to hours. Organizations can rapidly test, validate, and gain confidence through transparent structures that make AI reasoning inspectable and predictable. 2. **Time to Value:** While traditional deployments require six-month cycles, Amigo agents can be deployed in weeks. Our Agent Engineers facilitate an accelerated implementation journey that gives our partners a substantial head start in the market. 3. **Time to Flywheel**: Success ultimately depends on establishing a swift self-reinforcing improvement cycle, where data collection drives system enhancement that leads to wider adoption, thereby generating more data for further refinement. Our rapid iterative training approach creates this virtuous cycle by design. Our strategic advantage lies in helping our partners get to market safely *and quickly* so they can maintain their competitive advantage. The next 12 months represent a critical inflection point for organizations to establish their AI strategy in healthcare and beyond. Those who begin accumulating real-world AI interaction data now will secure decisive advantages as technology evolves. ### Join us in Building the Future At Amigo, we've built an agent architecture and platform that delivers uncompromising reliability without sacrificing implementation speed. We understand that in healthcare—where decisions impact lives—there is no room for error and no time to waste. We've been quietly building partnerships with some of the world's most forward-looking healthcare organizations. Today, our partners are building AI doctors, AI dietitians, and AI nurses to provide care coordination across complex patient journeys. In each case, our technology is not replacing healthcare professionals but amplifying their capabilities, maintaining a robust trust framework that validates every action. Soon, we'll be sharing these success stories and the measurable impact our approach is delivering in real-world healthcare environments. We invite forward-thinking leaders to join our partner program. Whether you're looking to enhance clinical decision support, streamline operations, or improve patient engagement, we provide a path to trusted AI implementation that respects the unique complexities of healthcare environments. [**Book a time**](https://cal.com/alikhokhar/30min?overlayCalendar=true) with me directly to learn more about what we’re building and how we can partner with you.