Skip to main content
3Nsofts logo3Nsofts
On-Device AI

A Native SwiftUI Bot That Feels Reliable: Streaming, Cancellation, and Recovery

The useful bot interface is more than a chat bubble: show availability, incremental progress, source-backed answers, cancellation, and a manual path when the model cannot respond.

By Ehsan Azish · 3NSOFTS··8 min read

It is easy to make a SwiftUI screen that looks like chat. It is harder to make one that handles a person waiting, changing their mind, losing model availability, or discovering that an answer was based on the wrong record. Those are the moments that decide whether a native bot is useful.

Apple's Foundation Models API can power on-device dialog on supported systems. The UI still needs a contract: what is being generated, what has been verified, and what the person can do if generation fails.

Start with the task, not a blank prompt

A blank “Ask anything” box gives the bot no useful scope and gives the person no clue about its limits. A stronger entry point starts from the current screen: “Explain this repair history,” “Find a date in my records,” or “Draft a summary of this note.” The selected record or collection becomes explicit context. The answer can then link back to the evidence and avoid pretending the model knows the whole library.

If you cannot name a task that is faster than normal search or filtering, a chat surface may not be the right feature. A native search view is often a better interface for exact titles, dates, and tags. Use language generation where interpretation, summarization, or a multi-step question actually adds value.

Show the states a person can act on

The conversation view should distinguish:

  • Unavailable: the model cannot run on this device, Apple Intelligence is off, or the model is not ready. Explain which case you know and offer a manual route.
  • Working: the request is running; show a stop control and keep the composer understandable.
  • Partial: streamed text may change or end with an error. Do not label it a final answer or save it as one.
  • Complete: the response finished; show source records and available follow-up actions.
  • Failed or cancelled: keep the person's question and any unsaved draft visible, with retry or edit.

Apple documents model availability states and session streaming. Check availability before presenting an action that depends on the model. Do not turn every unavailable state into a generic network error: an on-device model can be unavailable for reasons unrelated to connectivity.

Stream as a progress signal, not as proof

Streaming can reduce the feeling of waiting, but it can also make an incorrect answer look authoritative one word at a time. Render incremental output as provisional. Wait for the request to complete before enabling “copy final answer,” persistence, or a proposed write. If the bot is quoting saved records, validate that the referenced records still exist before showing a source chip.

The user can leave the view while generation is running. Give each request an identity and ignore late results from an older request after a newer one starts. Cancel the underlying task when the person taps Stop or when the workflow no longer needs it. Apple's cancellation and error guide is useful for the implementation boundary; the product rule is simpler: stopping generation must never leave an unannounced data change.

Use a transcript policy rather than retaining every turn indefinitely. Apple's LanguageModelSession maintains context between requests, so a long conversation can grow. Keep the task's relevant context, trim or start a new session when necessary, and separate private source data from user-visible conversation history. If the app stores chat history, disclose where it lives and make deletion work.

Make the answer inspectable

When a bot answers from the user's records, each factual claim should have a route back to its source. A result card can show the record title, date, and a button to open it. If no matching record was found, say that. Never transform “I did not find a warranty” into “you have no warranty”; the former is an observed search result, the latter is a broader claim.

For generative text that is not grounded in records, label it accordingly. If the app lets the bot propose an action, show the exact proposed change in a separate review step. The native agent action-boundary article explains how to keep model suggestions separate from commits.

Test the unhappy paths first

A polished animation cannot compensate for a bot that loses the person's question. Test these scenarios on a real supported device and with a deliberately unavailable model state:

  1. Ask a question, start another, and ensure the older response cannot overwrite the newer one.
  2. Cancel while text is streaming; verify no draft is saved as a completed answer.
  3. Remove a source record before the result is displayed; show a missing-source state.
  4. Deny an OS permission or make a local read fail; explain the limitation without inventing an answer.
  5. Use VoiceOver and Dynamic Type while a response arrives; avoid announcing every token and keep Stop reachable.
  6. Leave and return to the view; confirm the composer and completed history behave as promised.

Measure task completion, not merely number of messages. A useful bot helps someone find the right record or finish a draft, then gets out of the way. The underlying AI-native data architecture and Foundation Models versus Core ML decision guide are better starting points than a chat UI when the real requirement is classification or exact retrieval.

Authoritative References