Offline AI: Camera & Chat — my own iOS product, submitted to the App Store

Offline AI — An Assistant That Never Leaves the Phone

Offline AI — An Assistant That Never Leaves the Phone

Offline AI — An Assistant That Never Leaves the Phone

What changes when the model lives in your pocket?

What changes when the model lives in your pocket?

Client

Offline AI: Camera & Chat — my own iOS product, submitted to the App Store

My role

Sole designer and developer — product, UX, SwiftUI, on-device model integration

Timeline

2026

Impact

6 on-device language and vision models · 1,215 object labels · 246 seeded knowledge categories · works in airplane mode

The challenge

The challenge

A cloud assistant hides its constraints: memory, heat and latency are somebody else’s problem. On a phone they are the product. A model that fits in 230 MB is not as capable as one that needs 1.6 GB, the camera and the language model compete for the same chip, and a download that fails at 90% costs the user a gigabyte and their trust. I wanted an assistant that chats, sees, listens and talks with the network switched off, with no account and nothing sent anywhere. The question was not whether a small model can answer. It was what the interface has to do so that a small model, finite memory and a warm phone still feel like a capable assistant.

A cloud assistant hides its constraints: memory, heat and latency are somebody else’s problem. On a phone they are the product. A model that fits in 230 MB is not as capable as one that needs 1.6 GB, the camera and the language model compete for the same chip, and a download that fails at 90% costs the user a gigabyte and their trust. I wanted an assistant that chats, sees, listens and talks with the network switched off, with no account and nothing sent anywhere. The question was not whether a small model can answer. It was what the interface has to do so that a small model, finite memory and a warm phone still feel like a capable assistant.

Fig. 01 — One pipeline, one chipEverything on the phone
Sense
1/s
Camera, voice, text
Camera checked once a second; speech recognised on-device
Recognise
1,215
object labels
MobileCLIP matches each frame against the vocabulary
Remember
246
seeded categories
Knowledge graph that grows with what you teach it
Reason
LFM2
local language model
Streams tokens through Liquid AI's LEAP SDK
Respond
1
sentence at a time
Speaks while the reply generates, with Skip
No server anywhere in the loop.
The network is used once, to download a model. After that, airplane mode changes nothing.

What I did

What I did

1. Answered “will it fit?” before the download. The model manager reads the phone’s memory and free storage and labels every model Fits well, Heavy, Too large or No space before a byte is downloaded. iOS lets one app use roughly half the memory before shutting it down, so a model that technically loads can still be marked Heavy. People pick a model their phone can actually run.

1. Answered “will it fit?” before the download. The model manager reads the phone’s memory and free storage and labels every model Fits well, Heavy, Too large or No space before a byte is downloaded. iOS lets one app use roughly half the memory before shutting it down, so a model that technically loads can still be marked Heavy. People pick a model their phone can actually run.

Fig. 02 — Will it fit? Answered before the downloadInteractive · pick a phone
Memory
Free storage
ModelDownloadNeedsLabel
LFM 350M230 MB1 GB RAMFits well
LFM 700M450 MB2 GB RAMFits well
LFM Vision 450M380 MB2 GB RAMFits well
LFM Instruct 1.2B730 MB3 GB RAMHeavy
LFM Vision 1.6B1.1 GB4 GB RAMHeavy
LFM 2.6B1.6 GB5 GB RAMToo large

The same rule the app runs. iOS lets one app use roughly 55% of memory before it is shut down, so a model above that line is Heavy even if it loads. No space wins over everything: a download that cannot finish is never offered.

2. Designed when the agent speaks, not just what it says. Agent Vision checks the camera once a second and stays quiet until the same object holds for three seconds. Then it says one sentence of at most fifteen words, with Skip always on screen, and does not narrate the same object again for a minute. A narrator that talks constantly is a narrator people switch off.

2. Designed when the agent speaks, not just what it says. Agent Vision checks the camera once a second and stays quiet until the same object holds for three seconds. Then it says one sentence of at most fifteen words, with Skip always on screen, and does not narrate the same object again for a minute. A narrator that talks constantly is a narrator people switch off.

3. Gave vision and language their own turns. Recognition and the language model share one chip, and running both at once froze the camera feed in testing. Camera analysis now pauses while the model thinks and speaks, and resumes when the sentence ends. The scheduling became part of the interaction: the agent looks, then talks.

3. Gave vision and language their own turns. Recognition and the language model share one chip, and running both at once froze the camera feed in testing. Camera analysis now pauses while the model thinks and speaks, and resumes when the sentence ends. The scheduling became part of the interaction: the agent looks, then talks.

Fig. 03 — How Agent Vision decides to speakOne tick per second
Every second
Read the frame
Best label from your taught objects, else the vocabulary
Stable?
Same label 3 ticks in a row
Otherwise keep watching quietly
Taught by you →
Say
One sentence, 15 words max
Uses facts from the knowledge graph; Skip always on screen
Unfamiliar →
Ask
What should I call it?
By voice, once per object every 3 minutes
Learn
Name + visual signature
Stored on the phone, linked in the graph
While the model thinks and speaks, camera analysis pauses.
Vision and language share one chip. Running both froze the camera feed in testing, so they take turns. The same object is not narrated again for a minute.

4. Let people teach it. When an object is stable but unfamiliar, the agent asks by voice what to call it and listens for the answer. The name is stored on the phone with the object’s visual signature and a node in the knowledge graph, so next time it recognises this particular mug, not just “a mug”. It asks once, then leaves that object alone for three minutes.

4. Let people teach it. When an object is stable but unfamiliar, the agent asks by voice what to call it and listens for the answer. The name is stored on the phone with the object’s visual signature and a node in the knowledge graph, so next time it recognises this particular mug, not just “a mug”. It asks once, then leaves that object alone for three minutes.

The solution

The solution

A multimodal assistant designed around the device. Chat streams from a local LFM2 model through Liquid AI’s LEAP SDK. MobileCLIP matches camera frames against 1,215 labels. Agent Vision narrates and learns names, and a knowledge graph starts with 246 seeded categories and grows with what you teach it. The interface is native SwiftUI with system controls. After a model is downloaded, airplane mode changes nothing. The code is open source and version 1.0 is in App Store review, so there are no adoption numbers to report yet.

A multimodal assistant designed around the device. Chat streams from a local LFM2 model through Liquid AI’s LEAP SDK. MobileCLIP matches camera frames against 1,215 labels. Agent Vision narrates and learns names, and a knowledge graph starts with 246 seeded categories and grows with what you teach it. The interface is native SwiftUI with system controls. After a model is downloaded, airplane mode changes nothing. The code is open source and version 1.0 is in App Store review, so there are no adoption numbers to report yet.

Fig. 04 — The app, recorded on an iPhoneAirplane mode on

A local 1.2B model on the phone in airplane mode: suggested prompts, history, the model name in the title.

What I’d do differently

What I’d do differently

I started from what the models could do and treated the device as an engineering detail. Fit labels, one-sentence narration and turn-taking between vision and language all arrived later, as fixes for memory, latency and heat. They are the product, not patches. Next time I would put the device limits into the first sketch and test the whole flow on a 4 GB iPhone before designing anything for a Pro.

I started from what the models could do and treated the device as an engineering detail. Fit labels, one-sentence narration and turn-taking between vision and language all arrived later, as fixes for memory, latency and heat. They are the product, not patches. Next time I would put the device limits into the first sketch and test the whole flow on a 4 GB iPhone before designing anything for a Pro.

More work

All case studies →

All case studies →

All case studies →

d.trubnikov@me.com

© 2026 Dima Trubnikov · Vilnius, Lithuania (EU)