Beyond “It Works”: The PACE Framework for Teaching AI Coding Agents Reliability and UX

Coding Agent lesson learned

· coding,agent

AI coding agents are very good at producing code that works on the happy path. The harder skill is teaching them to ask: “What happens to the user when the network is slow, the server fails, or this side effect completes out of order?”

That requires more than telling an agent to “add try/catch.” A try/catch can hide failure just as easily as it can handle it. An await can preserve correctness—or add eight seconds of unnecessary user-facing latency.

Use the PACE framework to guide both humans and AI agents:

  • P — Product outcome
  • A — Await and ordering
  • C — Catch ownership
  • E — Evidence

The real problem: correct code can still create bad UX

Consider a request to clear an AI session’s memory after the AI fails to respond:

const result = await aiAgent.clearAgentMemory(session.sessionId);

This is syntactically correct. But it leaves important product questions unanswered:

  • Is memory clearing required before the user receives an error message?
  • Can the remote call take several seconds?
  • What happens on timeout, HTTP failure, or malformed streaming output?
  • Does retrying risk sending the reset command twice?
  • Could a background reset finish after the user sends their next message?

The question is not “Where is the try/catch?” The question is: which part of the system owns the failure, and what should the user experience be?

P — Product outcome

Before writing code, define what must be true for the user.

For example:

If the AI cannot respond, notify the user immediately. Attempt to clear the failed AI session in the background. A reset failure must be observable to operators but must not delay or prevent the user-facing fallback message.

This statement gives the agent a design target. It distinguishes the primary outcome—informing the user—from the secondary operation—cleaning up memory.

A stronger prompt is not:

Add a memory reset API call.

It is:

The user must receive the failure message immediately. Session cleanup is best effort and must not block the request. The cleanup operation must have a bounded timeout, no unsafe retries, and structured logging for failures.

A — Await and ordering

await is not inherently wrong. It means: “Do not continue until this operation settles.”

Use it when the next action would be incorrect without the result:

const payment = await chargeCustomer();await createOrder(payment.id);

Do not use it on a user-facing path when the operation is explicitly best effort:

void aiAgent.clearAgentMemory(sessionId).catch((error) => { console.warn('[ai-handlers] failed to clear AI chat memory', error);});

await sendUserFallbackMessage();

This lets the user receive the fallback message immediately while cleanup runs independently.

But “background” is not automatically safe. It creates an ordering question: what if the user sends another message before the memory reset completes?

An agent should identify that risk and recommend one of these choices:

  • Allow the next message immediately if stale memory is acceptable.
  • Mark the session as resetting and temporarily reject or queue new messages.
  • Await the reset if correctness is more important than response speed.
  • Send the user a clear status such as “Please retry in a few seconds.”

The correct choice is a product decision, not a coding-style preference.

C — Catch ownership

Do not teach agents to wrap every function in a broad try/catch. Teach them to assign responsibility.

A useful pattern is:

async function clearAgentMemory(sessionId) { try {
const result = await clearSessionOnce(sessionId);

if (!result.ok) {
console.warn('[qwenpaw] memory clear not confirmed', result);
}

return result;
} catch (error) {
console.warn('[qwenpaw] unexpected memory-clear failure', error);
return { ok: false, outcome: 'UNEXPECTED_ERROR' };
}
}

The helper owns transport and protocol failures. The caller owns the workflow decision: block, background, retry, notify, queue, or escalate.

Good error handling should answer:

  • What failed?
  • Was the action definitely not performed, or merely unconfirmed?
  • Is retry safe?
  • Who needs to know: the user, an operator, or both?
  • What information can be logged safely?

For a remote /clear command, a timeout may mean “unknown,” not “failed.” The server might still complete the reset after the client disconnects. Automatically retrying could create conflicting work or misleading logs.

E — Evidence

Reliability is proven by behavior, not by the presence of a try/catch.

Ask the agent to verify:

  • Successful response and expected confirmation marker
  • HTTP failure
  • Timeout and aborted request
  • Malformed or incomplete streaming response
  • No retry when retries are unsafe
  • User fallback message is sent immediately
  • Cleanup failure does not create an unhandled rejection
  • A new message arriving during cleanup follows the intended ordering policy

For example:

Add tests proving that a failed cleanup request does not delay the user fallback message, and that the cleanup request is issued only once.

A reusable instruction for AI coding agents

Put this in your repository’s AGENTS.md, custom instructions, or engineering playbook:

## Async Reliability and UX
Before changing asynchronous or error-handling code:

1. State the user-facing outcome and identify the critical path.
2. Decide whether each external operation is required, best effort, or deferred.
3. Explain whether `await` is necessary and what latency it can add.
4. Identify ordering and race risks if work is moved to the background.
5. Assign error ownership: helper, caller, queue, or user-facing boundary.
6. Do not add broad or silent catches. Log failures with safe context and return
structured outcomes where the caller must make a workflow decision.
7. For changes affecting latency, side effects, retries, or ordering, explain
the trade-off before implementation and add relevant tests.

A better prompt for day-to-day work

Use this when assigning a task:

Before implementing, perform a short PACE review: product outcome, await/latency implications, catch ownership, and verification plan. Highlight any user-experience or ordering trade-offs. Do not use retries or background execution unless their safety is established.

The goal is not to slow AI coding agents down. It is to direct their speed toward the right questions.

Great AI-assisted development is not “generate code, then patch failures.” It is teaching the agent to treat reliability, latency, observability, and user experience as part of the feature definition.