Local Character › Methodology · Last reviewed 2026-09-29
Methodology
Every claim this site makes about speed, size or privacy is listed here with how it was measured. Measurements were taken by the developer on an Apple Silicon Mac (Apple GPU) in Chromium with WebGPU, on the live site, in September 2026. Other hardware will differ.
The pipeline
| Job | Model | Runtime in the browser | Size | Licence |
|---|---|---|---|---|
| Chat replies | Gemma 4 E2B (instruction-tuned), gemma-4-E2B-it-web | LiteRT-LM web (@litert-lm/core 0.17.1) on WebGPU; needs WebAssembly JSPI and relaxed SIMD | 2,008,432,640 bytes | Gemma Terms of Use |
| Scene pictures | DreamShaper 8 with an LCM scheduler, 4 steps, 512×512 | ONNX Runtime Web on WebGPU | about 2.16 GB | CreativeML OpenRAIL-M |
| Spoken replies | Kokoro-82M (fp32), 28 English voices | ONNX Runtime Web on WebGPU; the final iSTFT runs in JavaScript | about 0.37 GB | Apache-2.0 |
| Speech input | Whisper base (fp32 encoder and merged decoder) | ONNX Runtime Web 1.27 on WebGPU, greedy decoding; log-mel features computed in JavaScript | about 0.29 GB | MIT |
Model files are served by SHA-256 from models.skillsafe.ai, SkillSafe's registry of vetted model files, and kept in the browser's Cache Storage after the first download. The ONNX models (pictures, voice and speech input) share one GPU queue, so they run one at a time instead of colliding.
Measured results
| What | Result | How | Date |
|---|---|---|---|
| Chat reply time | 6-8 s per reply in a 71-message chat | Live site, cached model | 2026-09-29 |
| Chat reply time, 35-turn scripted test | mean 8.1 s, slowest 13 s (the previous MediaPipe runtime: mean 17.8 s, slowest 39 s) | Same 35 turns replayed on both runtimes | 2026-09-28 |
| Time to first word | about 0.3 s, flat across turns | The conversation keeps its attention cache between turns | 2026-09-28 |
| Chat model ready after reload | about 4 s from the cached file | Live site | 2026-09-29 |
| Voice, cold start | about 28 s to the first sound | First reply after a reload; the app pre-warms the voice model to hide this | 2026-09-24 |
| Speech recognition accuracy | Token ids identical to Python ONNX Runtime on the same audio; log-mel features within 2×10-5 of the reference WhisperFeatureExtractor | Independent reference implementation, same input | 2026-09-24 |
| Picture model download | about 2 minutes for 2.2 GB | First picture on a fast connection | 2026-09-24 |
Weak results, stated plainly
- Small model, simple writing. Gemma 4 E2B has 2 billion parameters. It can lose track of details in very long scenes and sometimes repeats phrases from its own earlier replies.
- Narrow browser support. The chat runtime needs WebAssembly JSPI and relaxed SIMD, which only Chrome and Edge 137+ have today. Safari and Firefox get no chat.
- Half-precision voice was unusable. The fp16 Kokoro model ran fast on WebGPU but produced garbled speech (a transcriber heard "waiting" as "beaten"), so the app ships the larger fp32 model.
- Not an offline app. There is no service worker, so the page needs a connection to open, even though the models are cached and replies are generated on the device.
How privacy is enforced
The app has no chat endpoint to send messages to. Characters, chats and personas are written to the browser's IndexedDB; model files go to Cache Storage. The hosting platform sets a Content-Security-Policy that limits which hosts the page can contact. The privacy page lists every network request the app makes.