Skip to main content
3Nsofts logo3Nsofts
On-Device AI

Building a Native SwiftUI Agent: Let the Model Propose, Let the App Decide

A practical architecture for native iPhone agents: distinguish answers from actions, validate tool arguments, confirm consequential changes, and make every result observable.

By Ehsan Azish · 3NSOFTS··9 min read

A native agent becomes useful when it can do more than answer a question. It can find a record, prepare a change, or carry out a small workflow. That is also the moment an AI feature becomes an authority problem: who is allowed to decide that an action should happen?

SwiftUI provides the interface, and Apple's Foundation Models framework can generate text and call app-defined tools on eligible devices. Neither one makes the model a trustworthy owner of your data. The useful architecture is an agent that proposes, with ordinary app code retaining the right to authorize and commit.

Define three classes of request

Many bot prototypes expose every operation as a tool. That makes a request like “tidy my notes” dangerously ambiguous. Separate capabilities before writing a prompt:

| Request | Example | App behavior | | --- | --- | --- | | Read | “Which warranty expires next?” | Fetch records and answer with a source link. | | Prepare | “Draft a reminder for this receipt.” | Produce an editable draft; save nothing yet. | | Commit | “Delete the old receipt.” | Confirm exact target, validate access, then perform and record the result. |

The model may help interpret the request, but the app determines which class a tool belongs to. A read tool should never also mutate state as a side effect. A “prepare” tool should return a proposed object, not quietly persist one. The confirmation screen should state which records will change and what can be undone.

This is a product design decision as much as an engineering one. People will trust a bot that tells them what it did. They will stop trusting one that says “done” when a save failed, or that makes a plausible change to the wrong item.

Keep the model behind a narrow capability boundary

Apple's tool-calling documentation lets an app supply tools to a LanguageModelSession. The model can choose tools, and the app implements them. Newer generation options can require or disallow tool use for a request. This is useful for grounding an answer in current local data, but a required tool call is not proof that the requested operation is safe or correct.

An effective boundary has four checks:

  1. Validate arguments. An item identifier must resolve to a real item in the active account or library. Dates, quantities, and enum choices must satisfy app rules.
  2. Check permissions at execution time. A tool name in a session is not an authorization grant. OS permissions and the app's own access rules still apply.
  3. Confirm consequential writes. Present the target and proposed change in a native SwiftUI review view. Require an explicit user action to commit.
  4. Return an actual outcome. A tool reports success only after persistence succeeds. On failure, leave the draft intact and show a retry or manual path.

For a sensitive workflow, even a locally running model should receive only the minimum records needed to answer. “On-device” describes where inference happens; it does not make overbroad access to personal data a good design. Keep a clear distinction between data the user selected and data the agent inferred might be relevant.

The SwiftUI state machine should be visible

Represent a request as states such as idle → reading → proposing → awaitingConfirmation → committing → completed or failed. SwiftUI renders these states; the model does not control them. A cancellation during proposal should stop generation and leave stored data untouched. A cancellation during a committed write needs the same transaction and recovery rules as a non-AI action.

Show source records alongside an answer. A statement like “your annual service is due in October” should link to the saved service date that supports it. If the model could not find a record, say so rather than inferring one from general knowledge. If a result is a draft, label it as a draft and allow editing before save.

This architecture also gives you a clean test surface. Unit-test the action classifier and validator with ambiguous requests, deleted records, expired permissions, and duplicate taps. Integration-test the native confirmation view separately from the model. On device, test the full flow with the model unavailable as well as available: Apple says availability varies, and a useful manual path should survive.

When a chatbot is enough

Do not add agent behavior merely because the framework offers tools. If the user needs to search their documents and see a concise explanation, a read-only assistant can be simpler to understand and safer to ship. Add an action capability only when it removes real work, the target can be identified reliably, and the result can be reviewed or reversed.

For implementation details, read the Foundation Models tool-calling guide. For the broader data-model decision, see AI-native iOS architecture. The companion article covers the SwiftUI bot experience from streaming to recovery, and the iOS AI architecture audit lists production checks around inference and privacy.

Authoritative References