What is the difference between on-device AI and cloud AI?
Cloud AI sends your data to a company's servers, runs a large model there, and returns the result. On-device AI runs the model directly on your phone or computer, so the audio, document, or question is processed where it already lives. The practical difference is who can access your data, and what happens when there is no internet.
Most AI products are cloud products. ChatGPT, Otter, Fireflies, and ChatPDF all work the same way underneath: upload, process, download. That architecture exists for a reason, since very large models need very large hardware. It also means every transcript and contract you feed them passes through someone else's computer.
On-device AI inverts that. Apple Intelligence, and apps built on local models, run smaller neural networks that fit in a phone's memory. Nothing has to be uploaded for the AI to do its job.
Is on-device AI more private than cloud AI?
Yes, and the difference is structural rather than a matter of policy. A cloud service must receive your data to process it, so your privacy depends on the vendor's retention rules and security record. On-device AI removes the upload step, which means there is no server-side copy of your content to leak or subpoena.
Cloud vendors vary a lot here. Some train on customer content by default, some let you opt out, and some keep transcripts indefinitely to power search features. Reading those policies is real work, and they change without much notice. With local processing there is nothing to audit, because the model ran on your hardware and the output stayed there.
This matters most for content you may not be allowed to upload at all: interviews under NDA, patient conversations, client calls, unpublished research. For that material, "we encrypt data in transit" is not the same promise as "the data never left."
One honest caveat. "On-device" describes where the AI runs, not everything an app does. In Lyonesse, recording, transcription, summaries, speaker labels, document analysis, and Q&A all run 100% on-device, with no cloud AI or transcription provider behind them; optional iCloud sync exists and is off by default. When you evaluate any private AI app, confirm the AI pipeline itself is local, not just one feature of it.
Does on-device AI work offline?
Yes. Once a model is downloaded, it runs with no connection at all. You can transcribe a recording in airplane mode or summarize a PDF mid-flight. Meeting notes in a basement conference room with no signal work the same way. A cloud tool in the same situation shows a spinner and waits.
Offline also means predictable. A cloud round trip depends on your connection quality and the vendor's server load that minute. A local model is bound only by your chip, and Apple silicon has become genuinely quick at this class of work.
How can you tell if an app is really on-device?
Three quick checks settle it. Turn on airplane mode and try the core feature; a truly local app keeps working. Look for a model download step, because real local AI stores a model on your device and honest apps show its size. Then read the App Store privacy label for what gets collected.
Marketing language blurs this. Plenty of apps say "private" while sending audio to a cloud transcription API over an encrypted connection. Encryption in transit protects you from eavesdroppers, not from the company on the other end. The airplane-mode test cuts through the copy in ten seconds.
Is on-device AI accurate enough?
For speech to text in English, yes, and this is measured rather than assumed. In our published LibriSpeech benchmark, Apple's on-device SpeechAnalyzer scored 2.12% word error rate on clean speech, ahead of Whisper Small at 3.74%. Our second round found NVIDIA's Parakeet models in the same territory, and Lyonesse ships Parakeet TDT v3 as a downloadable engine. Those numbers cover English only; we have not benchmarked other languages.
Small local language models are a different story, and it helps to be honest about it. A 0.6B parameter model on your phone will not out-reason a frontier cloud model. What it does well are bounded jobs: summarizing a transcript, or answering questions about a document it can read. It can pull action items out of a meeting too. Those happen to be the tasks most people need an assistant for every day.
When is cloud AI still the better choice?
Cloud AI wins when a task needs frontier-scale reasoning or live knowledge of the web. It also wins at heavy generation, like producing images or long original drafts from scratch. A large hosted model will do that better than anything that fits on a phone today.
The point is not that one side wins everything. It is that a big share of everyday AI work (transcription, summaries, document Q&A) never needed a data center, and routing it through one is a privacy cost you do not have to pay.
On-device AI vs cloud AI at a glance
| Cloud AI (Otter, ChatPDF, cloud assistants) | On-device AI (Lyonesse) | |
|---|---|---|
| Where processing happens | Vendor servers | Your iPhone, iPad, or Mac |
| Works offline | No | Yes, after a one-time model download |
| Account required | Almost always | No |
| Server-side copy of your files | Yes, at least during processing | None; the AI pipeline runs locally |
| Used for model training | Depends on the vendor's terms | No, processing never reaches a server |
| Speed | Varies with connection and server load | Bound by your chip, consistent |
What can you actually run on-device today?
More than most people expect. Lyonesse is one concrete example of a full local AI stack on iPhone, iPad, and Mac:
- Transcription: Apple Speech built in, or a downloadable NVIDIA Parakeet v3 model (about 485MB) covering 25 European languages. An Auto setting picks the best engine available on your device.
- Summaries and chat: Apple Intelligence on supported devices, or a downloadable offline model (Qwen 3, about 400MB). Ask questions about one PDF or across your whole library.
- Speaker labels: on-device diarization works out who said what in a meeting recording, with no bot joining the call.
- Images: OCR reads text in photos locally, and an optional vision model (about 1.4GB) can describe what a picture actually shows.
No account is required, and the app is free to download. A few years ago this combination did not exist on a phone. Now it is a model download and a settings choice.
Related reading
Frequently Asked Questions
Is on-device AI more private than cloud AI?
Yes, structurally. A cloud service must receive your files to process them; on-device AI processes them on your own hardware, so no server-side copy of your audio or documents is created in the first place.
Does on-device AI need an internet connection?
Only to download a model the first time. After that, transcription, summaries, and document Q&A run with no connection at all. Airplane mode works fine.
Is on-device transcription as accurate as cloud transcription?
In our published English LibriSpeech benchmark, Apple's on-device engine scored 2.12% word error rate on clean speech and beat Whisper Small at 3.74%. On-device accuracy is competitive, at least in English, where we have measured it.
Can my phone really run an AI model?
Yes. The local models Lyonesse uses range from about 400MB to 1.4GB, sized for phone and laptop memory. Apple silicon runs them quickly, and the app only offers models your device can actually hold.
Get Lyonesse
Private, on-device transcription with AI summaries and cross-library Q&A. Works on iPhone, iPad, and Mac.
Get Lyonesse on the App Store