The Hardest Problem in AI Agents: Knowing When to Stop and Ask for Help

Why stopping criteria are the most underengineered component of AI agents, with patterns for confidence thresholds, escalation loops, and graceful degradation.

#ai-agents#stopping-criteria#reliability#production#safety
Cover image for the article: The Hardest Problem in AI Agents: Knowing When to Stop and Ask for Help

Every AI agent I have deployed to production has failed the same way at least once: it kept going when it should have stopped. Not a crash, not an error — just relentless, confident progression toward a wrong answer while burning tokens, hitting APIs, and occasionally mutating state that should not have been mutated.

The stopping problem in AI agents is not about token limits or timeout errors. It is about building systems that recognize their own uncertainty and choose inaction over confident wrongness. After shipping agents across three production systems handling 50,000+ daily requests, I have concluded that stopping criteria are the single most underengineered component in the agent ecosystem.

Why Agents Do Not Stop

Large language models have no native concept of "I should not proceed." Their training incentivizes completion. When you give an agent a goal and tools, the default behavior is to keep invoking tools until something resembles success or the context window fills up.

This creates three failure modes:

Failure ModeBehaviorConsequenceDetection Difficulty
Confident hallucinationAgent produces plausible but wrong resultSilent data corruptionVery hard
Thrashing loopAgent retries the same failing approachToken waste, API hammeringMedium
Scope creepAgent expands task beyond boundariesUnintended state changesHard
Graceful but wrongAgent returns a reasonable-looking fallbackDownstream system acts on bad dataVery hard

In our production system, 23% of agent failures in the first month were not errors — they were cases where the agent completed successfully with wrong results because it never recognized it was off track.

The Confidence Threshold Pattern

The first defense is explicit confidence scoring at each decision point. Instead of letting the agent choose tools freely, require it to score its confidence before and after each action.

from dataclasses import dataclass
from typing import Optional
import json

@dataclass
class AgentStep:
    action: str
    reasoning: str
    confidence_before: float  # 0.0 to 1.0
    confidence_after: Optional[float] = None
    result: Optional[str] = None
    should_continue: bool = True

class ConfidenceGatedAgent:
    def __init__(
        self,
        model_client,
        tools: list,
        confidence_floor: float = 0.6,
        confidence_drop_threshold: float = 0.2,
        max_steps: int = 10,
    ):
        self.model = model_client
        self.tools = tools
        self.confidence_floor = confidence_floor
        self.confidence_drop_threshold = confidence_drop_threshold
        self.max_steps = max_steps
        self.step_history: list[AgentStep] = []
    
    def execute(self, task: str) -> dict:
        for step_num in range(self.max_steps):
            # Ask the model to plan next action WITH confidence
            plan = self._plan_next_step(task, self.step_history)
            
            # Gate 1: Absolute confidence floor
            if plan.confidence_before < self.confidence_floor:
                return self._escalate(
                    reason="confidence_below_floor",
                    confidence=plan.confidence_before,
                    step=step_num,
                )
            
            # Gate 2: Confidence dropping between steps
            if self._detect_confidence_drop(plan):
                return self._escalate(
                    reason="confidence_declining",
                    confidence=plan.confidence_before,
                    step=step_num,
                )
            
            # Execute the action
            result = self._execute_action(plan)
            plan.result = result
            plan.confidence_after = self._assess_result_confidence(plan, result)
            
            # Gate 3: Post-action confidence check
            if plan.confidence_after < self.confidence_floor:
                return self._escalate(
                    reason="result_confidence_low",
                    confidence=plan.confidence_after,
                    step=step_num,
                )
            
            self.step_history.append(plan)
            
            if not plan.should_continue:
                return self._complete(plan)
        
        return self._escalate(
            reason="max_steps_reached",
            confidence=self.step_history[-1].confidence_after,
            step=self.max_steps,
        )
    
    def _detect_confidence_drop(self, current_plan: AgentStep) -> bool:
        if len(self.step_history) < 2:
            return False
        
        recent_confidences = [
            s.confidence_after for s in self.step_history[-3:]
            if s.confidence_after is not None
        ]
        
        if not recent_confidences:
            return False
        
        avg_recent = sum(recent_confidences) / len(recent_confidences)
        drop = avg_recent - current_plan.confidence_before
        
        return drop > self.confidence_drop_threshold
    
    def _escalate(self, reason: str, confidence: float, step: int) -> dict:
        return {
            "status": "escalated",
            "reason": reason,
            "confidence": confidence,
            "step": step,
            "partial_work": self.step_history,
            "recommended_action": "human_review",
        }

The Escalation Loop Architecture

Stopping alone is not enough. The agent needs to escalate to something — a human, a different agent, or a fallback system. The escalation architecture determines whether "stopping" means "failing gracefully" or "routing to resolution."

AI agent escalation loop architecture

class EscalationRouter:
    """Routes agent escalations to appropriate handlers."""
    
    def __init__(self):
        self.handlers = {
            "confidence_below_floor": self._handle_low_confidence,
            "confidence_declining": self._handle_declining_confidence,
            "result_confidence_low": self._handle_bad_result,
            "max_steps_reached": self._handle_stuck,
            "tool_error_repeated": self._handle_tool_failure,
        }
    
    def route(self, escalation: dict) -> dict:
        handler = self.handlers.get(
            escalation["reason"],
            self._handle_unknown
        )
        return handler(escalation)
    
    def _handle_low_confidence(self, esc: dict) -> dict:
        # Low confidence before action: agent does not know what to do
        if esc["step"] == 0:
            # Never started — task might be unclear
            return {
                "action": "clarify_with_user",
                "message": "Task unclear. Need clarification before proceeding.",
                "partial_work": None,
            }
        else:
            # Started but lost the thread
            return {
                "action": "human_review_partial",
                "message": f"Completed {esc['step']} steps but confidence dropped.",
                "partial_work": esc["partial_work"],
            }
    
    def _handle_stuck(self, esc: dict) -> dict:
        # Reached max steps without completion
        return {
            "action": "decompose_and_retry",
            "message": "Task may be too complex for single agent run.",
            "suggestion": "Break into subtasks or escalate to specialist agent.",
            "partial_work": esc["partial_work"],
        }

Production Metrics: Before and After Stopping Criteria

We deployed confidence-gated stopping criteria across our document processing agent (handling 12,000 documents daily) and measured the impact over 8 weeks:

MetricBefore (no stopping)After (confidence gating)Change
Silent wrong results23% of completions4.1% of completions-82%
Escalations to human0 (agents always "succeed")18% of runs+18%
Human effort per error45 min (discovery + fix)8 min (review escalation)-82%
Token cost per task4,200 avg2,800 avg-33%
End-to-end accuracy71%93%+31%
Mean time to correct answer12 min4 min-67%

The counterintuitive result: adding stopping criteria and escalation improved both accuracy AND reduced costs. Agents that stop early on uncertain paths burn fewer tokens than agents that complete confidently with wrong answers.

The Three-Signal Stopping Framework

After iterating through multiple approaches, we settled on a three-signal framework that catches failures at different stages:

Signal 1: Pre-Action Confidence (catches uncertainty) Before each tool invocation, the agent must articulate its confidence that this action will move toward the goal. Below threshold, stop and escalate.

Signal 2: Result Validation (catches wrong results) After each action, validate the result against expected patterns. Not just "did the API return 200" but "does this result make sense given the task context?"

class ResultValidator:
    def validate(self, action: str, result: str, context: dict) -> float:
        """Return confidence 0-1 that this result is valid."""
        
        checks = [
            self._check_format_expected(action, result),
            self._check_semantic_consistency(result, context),
            self._check_not_error_masquerading(result),
            self._check_magnitude_reasonable(result, context),
        ]
        
        # Weighted average with veto: any check below 0.3 forces escalation
        if any(c < 0.3 for c in checks):
            return 0.0
        
        weights = [0.2, 0.4, 0.2, 0.2]
        return sum(c * w for c, w in zip(checks, weights))

Signal 3: Progress Trajectory (catches loops) Track whether the agent is making progress toward completion. If the last 3 actions did not measurably advance the task state, the agent is likely stuck in a loop.

def detect_thrashing(step_history: list[AgentStep], window: int = 3) -> bool:
    if len(step_history) < window:
        return False
    
    recent = step_history[-window:]
    
    # Check for repeated actions
    actions = [s.action for s in recent]
    if len(set(actions)) == 1:
        return True  # Same action repeated
    
    # Check for oscillating confidence
    confidences = [s.confidence_after for s in recent if s.confidence_after]
    if confidences and max(confidences) - min(confidences) < 0.05:
        # Confidence flat = no progress
        return True
    
    return False

Cost of Not Stopping: A Case Study

In March 2026, one of our customer support agents hit a case it could not resolve: a refund request that required accessing a system the agent did not have tools for. Without stopping criteria, here is what happened:

  1. Agent tried the refund tool (failed, no access)
  2. Agent tried to find an alternative approach (searched knowledge base)
  3. Agent composed a response explaining the refund was processed (hallucinated)
  4. Agent sent the message to the customer
  5. Customer waited 3 days for a refund that was never initiated
  6. Customer filed a complaint

Token cost of this failure: $0.12 Business cost of this failure: $2,400 (refund + compensation + escalation handling + trust damage)

The ratio of token cost to business damage was 1:20,000. Spending $0.02 on an escalation check would have prevented $2,400 in damage. The economics of stopping criteria are absurdly favorable.

Implementation Checklist

For teams implementing stopping criteria in production agents:

  1. Establish a confidence floor per domain. Start at 0.6, tune based on false escalation rate. You want 10-20% escalation rate, not 50%.
  2. Instrument every tool call with pre/post confidence. This is your observability layer for agent decision quality.
  3. Build escalation routing, not just stopping. An agent that stops without routing creates a dead end.
  4. Track the confidence-accuracy correlation. Calibrate your model — when it says 0.8 confidence, is it actually right 80% of the time?
  5. Design for graceful partial completion. The agent should save partial work so the human reviewer starts from progress, not from zero.

Key Takeaways

  1. LLMs default to completion, not correctness. You must engineer stopping behavior explicitly because the model will always produce an output.
  2. 23% of agent "successes" were silent failures before we added confidence gating. Your success rate is almost certainly overstated.
  3. Stopping criteria reduce cost by 33%. Agents that bail early on wrong paths spend fewer tokens than agents that complete confidently with wrong answers.
  4. The escalation path matters more than the stop signal. Stopping without routing is just a polite crash.
  5. Calibrate confidence thresholds empirically. Start conservative (high escalation rate), then lower the floor as you gain trust in the model's self-assessment accuracy.

The mature AI agent is not the one that always completes its task. It is the one that knows the boundary between its competence and its ignorance — and acts on that knowledge before causing harm.

Comments

    No comments yet. Be the first to share your thoughts.