Real-time AI (audio/video in, voice out) on an M3 Pro with Gemma E2B

Posted by ffinzy@reddit | LocalLLaMA | View on Reddit | 32 comments

Sure you can't do agentic coding with the Gemma 4 E2B, but this model is a game-changer for people learning a new language.

Imagine a few years from now that people can run this locally on their phones. They can point their camera at objects and talk about them. And this model is multi-lingual, so people can always fallback to their native language if they want. This is essentially what OpenAI demoed a few years ago.

Repo: https://github.com/fikrikarim/parlor