Skip to content

Local AI Is No Longer a Compromise. It's a Design Decision.

#local-ai #privacy #on-device-inference #data-sovereignty #gdpr #pii-redaction

Local AI Is No Longer a Compromise. It's a Design Decision. ​

The cloud API habit is a dependency you chose ​

Every AI feature you've shipped this year is probably a cloud call. Paste text into an API, wait for a server farm in Virginia, parse the JSON. That pattern has become so normal that most developers don't question it. Cyrus Lopez, who built the Brutalist Report news aggregator, argues this is a design failure, not a convenience.

The point isn't that cloud models are useless. It's that the default is backwards. A feature that could run on the device in your pocket now depends on network conditions, vendor uptime, rate limits, account billing, and your own backend health. In his words: you took a UX feature and turned it into a distributed system that costs you money. The moment you stream user content to a third-party provider, you also inherit data retention questions, consent obligations, audit trails, breach notifications, and government requests. A summary feature now has a legal department.

And the hardware is sitting there idle. The Neural Engine in a modern phone crunches ML workloads at a pace that would have filled a server room a decade ago. Most of the time it waits for a JSON response.

"AI everywhere" is the wrong goal. Useful software is the goal. If the feature can be done locally, shipping it through a cloud API is self-inflicted damage.

There are use cases that demand cloud intelligence. But those aren't most use cases. Summarize, classify, extract, rewrite, normalize: that's the work most AI features actually do, and it's exactly the work local models have become excellent at. A local model doesn't need to write Shakespeare. It needs to summarize the page you just loaded, and for that, it's fast, free, and private. Lopez's closing line lands hard: you don't build trust with a 2,000-word privacy policy. You build trust by not needing one.

What local-first actually looks like ​

The Brutalist Report's iOS client is a clean example. The summaries run on-device through Apple's local model APIs. No server detours, no prompt or user logs, no vendor account, no "we store your content for 30 days" footnote. The input is already on the device because the user is reading the article there. The output is a lightweight Markdown summary. Cloud intelligence would add nothing but risk.

swift
import FoundationModels

let model = SystemLanguageModel.default
guard model.availability == .available else { return }

let session = LanguageModelSession {
    """
    Provide a brutalist, information-dense summary in Markdown format.
    - Use **bold** for key concepts.
    - Bullet points for facts.
    - No fluff. Just facts.
    """
}

let response = try await session.respond(
    options: .init(maximumResponseTokens: 1_000)
) {
    articleText
}
let markdown = response.content

For longer articles, chunk the text at roughly 10k characters, produce facts-only notes per chunk, then run a second pass to merge them into a final summary. This is the kind of task local models are perfect for: transforming user-owned data, not acting as a search engine for the universe.

The newer pattern is even better. Instead of "ask the model for JSON and pray," you define a struct and let the model fill it in. Your UI gets typed fields instead of scraped Markdown, and the model becomes a trustworthy subsystem instead of a novelty.

swift
@Generable
struct ArticleIntel {
    @Guide(description: "One sentence. No hype.") var tldr: String
    @Guide(description: "3-7 bullets. Facts only.") var bullets: [String]
    @Guide(description: "Comma-separated keywords.") var keywords: [String]
}

let session = LanguageModelSession()
let response = try await session.respond(
    to: "Extract structured notes from the article.",
    generating: ArticleIntel.self
) {
    articleText
}
let intel = response.content

Three architectures are in play now, and they behave very differently:

The third path is newer and quietly important. More on that later.

The hardware reality check ​

CanIRun.ai is the tool I wish existed two years ago. It reads your GPU or your Mac's unified memory from the browser, matches it against a hardware database, and ranks the open models that actually fit. Not just chat models. Image generation, video, reasoning, coding. 104 models across 302 devices from 22 labs, scored on speed, memory headroom, and quality.

When I ran it on my own laptop, the ranking confirmed what I suspected: the best fit wasn't the biggest model. It was the one that left headroom for the context window.

The memory guide is refreshingly honest:

MemoryWhat usually fitsExample device
8 GB3B–9B chat models at Q4, plus tiny image modelsRTX 4060
12 GB12B–14B dense models, smaller MoERTX 4070
16 GB20B–27B at Q4, comfortable 14B at higher qualityApple M4
24 GB30B dense, mid-size MoE, or local videoRTX 4090
32 GB+Larger open MoE and high-quality image generationRTX 5090

That 8 GB row is the one to internalize. A 7B–9B chat model at Q4_K_M quantization fits on an RTX 4060, which means a mainstream gaming laptop can serve a competent local assistant. The quant matters: Q4_K_M is the sweet spot for local chat, much smaller than full precision with only a modest quality drop. Q8 and F16 look closer to the original but eat VRAM for marginal gains.

One trap the site calls out: mixture-of-experts models still load all experts into memory even when only a few are active per token. A 47B MoE with 12.9B active isn't a 13B model for memory purposes. It's a 47B model. That's the most common hardware mistake I see people make.

Quick Take: Local models are now good enough for most production features. The remaining fight is about consent, not capability.

The server-grade local stack ​

The "local models are toys" objection died when TensorFold shipped. It serves serious text models on Apple Silicon and NVIDIA GPUs through an OpenAI-compatible API. You point your client at http://127.0.0.1:8080/v1 and existing code keeps working. No provider SDK migration. Just a different base URL.

The interesting part is speculative decoding. TensorFold supports multi-token prediction (MTP) heads and draft models like DFlash2, with a strict guardrail: a draft is accepted only when it equals the token the same engine would produce serially. That exactness guarantee matters for reproducible workloads, and it's rare in local serving.

ModelBackendDrafting
Nemotron 3.5 Lightning (30B-A3B)MLX, CUDAIncluded MTP head
Qwen3.8-27BMLX, CUDADFlash2 + context copies
Qwen3.8 Flash NextMLX, CUDAIncluded MTP head
GLM-5.3-FlashMLX (256 GB Mac), CUDA 2 ranksMTP, optional DFlash2
Gemma 4 26B-A4BMLXContext copies, optional DFlash2

When I tested a similar setup, the surprise wasn't the speed. It was how boring the integration was. Same API, same streaming, same chat shape. Local serving has quietly matured into infrastructure.

The heavy rows are honest about their appetites. GLM-5.3-Flash wants a 256 GB Mac or a two-rank CUDA setup. On CUDA, the auto setting serves one request at a time; you have to explicitly opt into shared rounds. If you need concurrency, budget accordingly. The fact that these trade-offs are documented in release notes with measured numbers puts TensorFold ahead of most commercial serving layers I've used.

The dark side: Chrome's silent 4 GB model ​

Local AI has a consent problem, and Google just demonstrated it at planetary scale. Chrome is writing a 4 GB on-device AI model, the weights for Gemini Nano, into user profiles without asking. No dialog. No settings checkbox. The file lands in a directory named OptGuideOnDeviceModel and powers features like "Help me write" and on-device scam detection, which are enabled by default on eligible hardware.

The investigation that broke this story is a masterclass in digital forensics. A freshly created Chrome profile with zero human input: the audit driver only loaded pages via the DevTools protocol and never touched the omnibox. The macOS filesystem event log recorded the entire sequence. Directory created at 16:38:54. Three unpacker subprocesses spawn at 16:47:22. The final move into place at 16:53:22. Total install time: 14 minutes and 28 seconds, all while a tab sat idle waiting for a five-minute timer.

4 GB of Gemini Nano weights silently written to disk, with zero user input. Install time: 14 minutes and 28 seconds, triggered during an idle tab. The deletion loop: delete weights.bin, Chrome re-downloads it on the next eligible window. The only fixes: chrome://flags, enterprise policy, or uninstalling Chrome. The climate bill for one model push at Chrome's scale: 6,000 to 60,000 tonnes of CO2-equivalent emissions. Chrome's market share: above 64%, a user base of 3.45 to 3.83 billion people.

This one hit close to home. I caught my own Windows install filling up last year and spent an afternoon chasing a 4 GB file under AppData I didn't recognize. I deleted it. Two weeks later it was back. Community reports of the same loop go back over a year, and the pattern maps almost exactly onto Anthropic's silent Native Messaging bridge install, which the same investigator called out two weeks earlier: forced bundling across trust boundaries, invisible defaults, harder to remove than install, automatic re-install on every run.

The legal exposure is severe. The investigation argues this breaches Article 5(3) of the ePrivacy Directive, which requires consent before storing information on a user's device, plus the GDPR's lawfulness, fairness, and transparency principles and its data-protection-by-design obligation. The environmental framing adds a new dimension: at Chrome's scale, one unrequested model push costs the planet thousands of tonnes of CO2 before a single user invokes an AI feature.

The lesson is uncomfortable: on-device AI is a privacy improvement only when the user opts in. Local execution removes the data-transfer problem, but it doesn't remove the consent problem. Chrome's architecture proves that a model can sit entirely on your disk and still be a violation. The decentralization that protects your data can also be used to ignore your choices.

The middle path: anonymize locally, reason in the cloud ​

Some tasks need frontier-model intelligence, and some data is legally protected. A law firm summarizing a contract has no interest in a 9B model's guess at legal reasoning. But pasting a client's contract into ChatGPT is a GDPR failure that can end careers.

The hardware math is brutal in the other direction. A frontier-grade open model needs 9,000 to 10,000 euros of hardware to serve. The small models that fit a normal laptop aren't in the same league on dense legal documents. So rizzo-pii takes a third route: keep the frontier model, remove the data from the equation.

The workflow is simple. Locally, a 0.3B model tags every span of personal data and replaces each one with a stable placeholder: [FULLNAME_1], [IBAN_1], [CF_1]. The mapping from placeholder to real value stays in a local dictionary. Identical values share the same placeholder, so the frontier model still sees coherent text and can reason about it. The anonymized text goes to ChatGPT, Claude, or Gemini. The answer comes back, and a local pass swaps the placeholders for the true values. The provider never receives a single real name, code, or number.

The model is the privacy layer, and it's cheap by design. 0.3B parameters on an mmBERT backbone runs on a CPU in 0.5 to 1.2 GB of RAM, quantized. No GPU. No API key. No telemetry. The privacy guarantee is structural rather than a promise: the data never leaves the device except as placeholder text.

The Italian-legal coverage is what sets it apart. Codice fiscale, partita IVA, dati catastali. Generic English-first models don't have labels for the three most sensitive identifiers in an Italian legal sentence, so they leave them in the clear.

Propertyrizzo-pii:0.3BOpenAI Privacy FilterMS Presidio
TypeDense encoder (mmBERT)Sparse MoE encoderNER + rules pipeline
Parameters~0.3B dense1.5B total / ~50M activespaCy + rules
Memory to load0.5–1.2 GB1.5B params residentVaries (spaCy)
Runs onCPU, under 1 GB RAMOn-deviceCPU
Categories22 (incl. IT-legal)8 genericConfigurable, EN defaults
Italian CF / PIVA / catastoYesNoNot by default
Checksum validationYesNoSome recognizers
Reversible mappingYes (local dict)MaskingAnonymization

The checksum layer is the quietly brilliant part. IBANs, VAT numbers, and credit cards must pass mod-97 or Luhn validation, and a valid checksum overrides the neural model. That eliminates the classic failure mode of a neural tagger fragmenting a long code, and it gives mathematically certain detection for the identifiers whose leakage is most damaging. On a 7,000-row held-out real Italian benchmark, the model hits 0.987 micro F1 and 0.998 token accuracy, with all five Italian-legal tags scoring a perfect 1.000.

For an architecture that promises "frontier models without giving up your data," those numbers are the difference between trust theater and a working system. Law firms, accountants, and notaries are the obvious customers, but the pattern generalizes: any regulated industry that wants frontier intelligence without the data-transfer liability.

Common pitfalls ​

After reading through these projects and testing some of them, here's what trips people up.

Sizing the model for the wrong job. A 70B dense model is overkill for extracting action items from notes. The canirun.ai guide keeps saying the same thing: 7B–9B at Q4 fits 8 GB and handles most chat and extraction tasks. Use local models as data transformers, not replacements for the entire internet, and you'll rarely need the big iron.

Forgetting that MoE models load all experts. The 47B MoE with 12.9B active is still a 47B memory footprint. People spec hardware by active parameters and then wonder why swapping fails. The weights are resident. All of them.

Treating quantization as a blunt instrument. Q4_K_M exists for a reason. fp16 quadruples the weight footprint and won't fit on the 8 GB card that handled the Q4 file fine. Going below Q4 to save another gigabyte degrades output quality fast. Match the quant to the task, not to the download size.

Trusting the privacy of a "local" model you didn't install. Deleting Chrome's weights.bin treats the symptom. Chrome's feature flags still mark the profile eligible, so the variations server schedules another push. If you want the model gone, you have to disable the feature, not delete the file. The same lesson applies to any runtime that re-installs its own components: the installer is the consent mechanism, and you need to know which one you have.

Sending raw documents because the vendor has a privacy policy. A privacy policy is not a legal basis for processing. The rizzo