Insights / AI Operations · · 11 min read
Designing human escalation paths in AI products
Every AI product needs a clear point where the machine stops and a person takes over. How we decide what gets escalated, how the handover works without losing context or leaking data, and how KeşifAtlası, CastLyra and EduRelia each design escalation for very different kinds of risk.
Every AI product reaches questions it should not answer on its own. Some are too consequential. Some fall outside the rules it was built on. Some involve a person who is upset, confused or at risk. The question is not whether those moments happen, but whether the product was designed for them.
Human escalation in AI products is that design: deciding in advance when the AI stops and a person takes over, how the handover works, what the user is told, and how each case makes the system better. It is not an afterthought or a "contact us" link at the bottom of the page. At Oryvelon, escalation is written into each company's source of truth before the AI feature is built. This article describes how we approach it, with three companies that face very different kinds of risk: KeşifAtlası, CastLyra and EduRelia.
Why escalation belongs in the design, not the support queue
Teams often treat escalation as a support problem. The AI handles what it can; anything else ends up in a general inbox. That arrangement fails in predictable ways.
The AI does not know it should stop, so it attempts answers to questions it should have passed on. When a case does reach a person, it arrives without context, and the user has to explain everything again. Nobody has decided who is responsible for which kind of case, so urgent ones wait behind routine ones. And because escalations are not tracked as a category, the product never learns from them.
Designed escalation fixes each of those. It gives the AI explicit reasons to stop. It hands the person a structured case. It assigns ownership. And it treats every escalation as information about where the product, the rules or the prompts need work.
It also follows directly from how we build AI products. If AI explains, verified data decides, then there must be a defined answer to "what happens when verified data does not cover this case?" Escalation is that answer.
What triggers an escalation
We define triggers per product, in writing, and review them at each 30/60/90-day checkpoint. They fall into five groups.
1. Consequential outcomes. Some decisions matter too much to leave to an automated explanation alone: a result that could change whether someone applies for a visa, a verification decision that affects someone's ability to work, a concern about a child. For these, a person reviews before or alongside the automated result, depending on the product.
2. Gaps in the rules or data. The verified rules database has no rule for this combination of circumstances. The store data is incomplete. The curriculum does not cover the question. The AI must not fill a gap with a plausible guess.
3. Disputes. The user disagrees with an outcome, or reports that something the product showed them is wrong. Disputes always go to a person.
4. Explicit requests. The user asks for a person. We honour that without making them argue with the AI first.
5. Safety and wellbeing signals. Language suggesting distress, a risk to someone, a child in a situation that needs an adult, harassment on a marketplace. These go to people quickly, with the product's safety procedures, and the AI's role becomes limited to acknowledging and routing. See Safety boundaries for consumer AI products.
Why model confidence is not enough
It is tempting to trigger escalation mainly on model confidence: if the AI is unsure, hand over. Confidence signals are useful, but they are not sufficient. Models can be fluent and wrong, and a confident answer to a question the product should never answer is worse than an uncertain one.
So most of our triggers are about the situation, not the model. A disputed eligibility result goes to a person whether the AI sounded confident or not. A question about a rule that is not in the database goes to a person even if the model could write a convincing paragraph. Confidence is one input among several, and it is never allowed to override a situational trigger.
How a good handover works
A handover has three audiences: the user, the person receiving the case, and the system that will learn from it.
For the user, the product says clearly that a person will look at this, what they can expect and roughly when. It does not pretend the handover is instant if it is not. It does not ask the user to repeat what they have already said.
For the reviewer, the case arrives structured:
- the user's question or the situation, in their own words where appropriate;
- the verified facts the product already has for this case;
- the rules or data that applied, and where the gap or conflict is;
- what the AI said, if anything, clearly marked as AI output;
- the trigger that caused the escalation;
- a priority, based on the trigger type.
The summary is built largely from structured data rather than generated prose, for the same reasons we validate AI outputs everywhere else. See Structured outputs: validating AI responses before users see them. Where the AI does write a short summary for the reviewer, it is labelled as such, and the underlying facts sit beside it.
For the system, every escalation is logged with its trigger, its outcome and what, if anything, should change: a new rule, a prompt adjustment, a new test case, a clearer onboarding message.
Escalation without data leakage
A handover moves information to a person, so it has to follow the same data rules as the rest of the product.
The reviewer sees what the case needs and nothing more. Access runs through a defined role with least-privilege access and two-factor authentication. The case stays inside the product where it started: a KeşifAtlası case is handled in KeşifAtlası's systems, by people working for KeşifAtlası, and never enters a group-wide queue alongside other companies' users. There is no cross-product user data, and escalation is not an exception to that. See Shared infrastructure, separate data.
Retention follows the product's rules. Escalated cases can contain sensitive details, so they have defined retention periods and deletion processes.
KeşifAtlası: when the rules run out
KeşifAtlası helps people understand visa and relocation eligibility, starting from Türkiye. Eligibility results come from a verified rules database. The AI explains which rules applied and why. It never decides.
Its escalation design follows from that structure.
Rule gaps. Imagine a user whose circumstances combine two situations the rules database does not address together — say, a change of employment status partway through a qualifying period. The rules engine returns "not covered" rather than a result. The AI does not attempt an answer. The user sees a clear message that their case needs a person, and the case goes to a reviewer with the answers they gave, the rules that were evaluated and the specific point where coverage ran out.
Disputes. A user believes their result is wrong. Perhaps they have read something on an official site that seems to contradict it. The dispute goes to a person who checks the rule, its source and its effective date. If the rule was outdated, the rules team updates the database and the case becomes a test case for the explanation prompt. See Prompt versioning and evaluation for production AI.
Paid reports. For the paid report, a person is part of the process for the harder questions by design. The AI prepares the explanation; people review the cases the product marks as complex. That is a large part of what the user is paying for, and the product says so.
Throughout, the framing matters. KeşifAtlası explains rules and routes hard questions to people. It does not present itself as legal advice, and escalation messages are written to keep that line clear. See Rules engines for eligibility.
CastLyra: when trust between people is at stake
CastLyra is a marketplace connecting verified creative talent — models, creators, creative professionals — with brands. AI helps structure profiles and briefs. The consequential moments are about people, not rules.
Verification decisions. Checks can be partly automated, but a decision that would reject or limit a talent's profile is reviewed by a person. Automated checks can flag; they do not have the final word. See Verification in talent marketplaces.
Consent and contact details. CastLyra has strict rules about when a brand can see a talent's contact details. If a brand's message tries to get around those rules — asking a talent to share a phone number before the platform's conditions are met, for example — the case is flagged for a person, not handled by an automated warning alone.
Safety concerns. Reports of harassment, misleading briefs or pressure to work outside agreed terms go straight to people with the authority to act on accounts. The AI's role is limited to acknowledging the report and routing it. It does not assess the credibility of the person reporting.
Profile disputes. If talent disagrees with how the AI summarised their profile, they edit it directly, and the edit wins. If a brand disputes a profile's accuracy, a person reviews it against the verified data.
On a marketplace, trust is the product. A slow or opaque escalation path on a safety report does more damage than any number of good matches can repair.
EduRelia: when a child is involved
EduRelia provides AI learning support for students, families and schools, with strict child data minimisation. Escalation here has a particular shape, because the right "person" is often not a support agent but a responsible adult already in the child's life.
Learning questions beyond scope. If a student asks about something outside the lesson's scope, or the assistant cannot give a reliable explanation, it says so and points the student to their teacher. It does not improvise.
Wellbeing signals. If a student's messages suggest distress or a situation that needs an adult, the assistant stops being a tutor. It responds with care, directs the student to a trusted adult, and the school's agreed procedure applies. What happens next is defined in the school's licence agreement, not decided by the product on the fly.
Content concerns. A teacher or parent who believes the assistant said something inappropriate can report it. The report goes to a person at EduRelia, the relevant output is reviewed, and any fix goes through the evaluation process before release.
Minimal handover data. Because of the product's data minimisation rules, handovers carry as little as possible: the relevant exchange and the minimum context, not a profile of the student. See Data minimisation for children in edtech.
What the user sees
Across all three products, the escalation message follows the same pattern:
- What is happening. "A person will review this."
- Why. "Your situation isn't covered by the rules we can check automatically." Or: "You asked to speak to someone." Short and honest.
- What to expect. A realistic time frame and how they will be contacted, within the product.
- What they can do meanwhile. Where relevant, what information would help, or where to find official sources.
We avoid two things in particular: pretending a person is involved when they are not, and making escalation feel like a failure on the user's part. Being passed to a person should feel like the product working as intended — because it is. Making the route to a person visible from the start is part of how we build trust during onboarding. See Earning user trust in AI products.
Staffing escalation in a small company
Escalation needs people, and our companies are designed to be lean. Three practices keep that manageable.
Priorities by trigger. Safety signals and disputes on consequential outcomes are handled first. Rule gaps are handled in order. Explicit requests for a person get a prompt acknowledgement even when the answer takes longer.
Escalation budgets. Each product tracks its escalation rate. A rising rate is a signal to fix something upstream — a missing rule, an unclear onboarding message, a prompt that attempts too much — rather than simply adding reviewers.
Degraded modes that include people. When an AI feature is unavailable, some products can route affected requests to people for a period instead of failing entirely. That choice is made per product, and the capacity limits are written down.
Learning from every escalation
An escalation is a small, expensive event. It is worth extracting everything from it.
Each case is tagged with its trigger and outcome. Monthly, each company's owner looks at the patterns as part of the regular monthly review:
- Rule gaps that recur become new rules in the database.
- Disputes that turn out to be right become fixes and test cases.
- Questions the AI attempted but should not have become new triggers.
- Confusion that stems from onboarding becomes clearer copy.
Over time, the aim is not zero escalations. It is fewer avoidable ones, and fast, well-handled ones for everything that genuinely needs a person.
Writing the escalation section before launch
Every Oryvelon company's source of truth has a short escalation section, written before the first AI feature ships. It answers five questions:
- Which situations always go to a person? Listed by trigger type, with examples specific to the product.
- Who receives each type of case? A named role, not "the team".
- What does the reviewer see? The fields in the structured case, and anything explicitly excluded.
- What is the user told, and what response time is realistic? Written as the actual message copy, not a note to write it later.
- How does each case feed back? Which log, which review, which owner.
Writing this down early changes the product. Teams discover triggers they had not considered, such as what happens when a user asks the same question three times in slightly different words, hoping for a different answer. They notice data the reviewer would need that the product does not currently keep, or data the product keeps that no reviewer should see. And they set response-time promises they can actually meet with the people they have, rather than ones that sound good on a landing page.
The section is short, usually less than a page. It is reviewed at each 30/60/90-day checkpoint along with the escalation numbers, and it changes whenever the product's scope does.
Common mistakes
- No written triggers. The AI keeps answering because nobody told it when to stop.
- Confidence-only escalation. Fluent, confident, wrong answers slip through.
- Context lost at handover. The user repeats everything, and the reviewer starts from zero.
- Hiding the route to a person. Users who cannot find it lose trust in everything else.
- Group-wide queues. Mixing cases from different companies breaks data boundaries.
- Escalation as a dead end. Cases are closed but never feed back into rules, prompts or tests.
Summary
Human escalation is the designed point where an AI product stops and a person takes over. We define triggers in writing for each product — consequential outcomes, gaps in rules or data, disputes, explicit requests and safety signals — and never let model confidence override them. Handovers carry a structured case with verified facts so users do not repeat themselves, stay inside the product that owns the data, and tell the user honestly what happens next. KeşifAtlası routes rule gaps and disputes to people, CastLyra keeps verification and safety decisions with people, and EduRelia hands wellbeing concerns to the responsible adults. Every escalation feeds back into rules, prompts and test sets, so the product gets better at knowing when to ask for help.
Questions and answers
What is human escalation in an AI product?
Human escalation is the designed handover from an AI system to a person when a question is too consequential, uncertain, disputed or sensitive for the AI to handle, including how context is passed and what the user is told.
When should an AI hand over to a human?
When the answer would decide something important for the user, when verified rules or data do not cover the case, when the user disputes an outcome or asks for a person, and when there are safety or wellbeing concerns.
How do you escalate without exposing personal data?
Pass the reviewer only what the case needs — a structured summary, the verified facts and the relevant rules — through a role with least-privilege access, and keep the case inside the product where it originated.