LemoLM: Local LLM & AI Chat
LemoLM: Local LLM & AI Chat is #27 in Developer Tools Grossing in Italy and charting in 1 of 24 countries we track.
Updated 27 Sep 2026 · rankings refresh through the day
DeveloperPAUL ANTHONY HADFIELD
CategoryDeveloper Tools
PriceFree
Released
Rating★ 0.0 (0)
Version1.7.0
Age rating12+
First seen by us
25 Sep 2026
Rating
0.0 ★
0 ratings
Measured
Best rank now
#27
Measured
LemoLM: Local LLM & AI Chat on the App Store
App details
App ID
6768808566
Publisher
PAUL ANTHONY HADFIELD
Content rating
12+
Where it ranks now
Every top chart it appears in across 24 countries.
| Country | Charts | Rank |
|---|---|---|
| Developer ToolsGrossing | #27 |
Apps near LemoLM: Local LLM & AI Chat on the chart
Developer Tools Grossing, around #27.
Ratings
App Store, worldwide.
Rating
0.0 ★
0 ratings
Measured
Ratings
0
Measured
About LemoLM: Local LLM & AI Chat
Explore Qwen3, DeepSeek R1, Mistral, Phi-4 and other local models directly on your iPhone or iPad.
On a suitable Apple silicon Mac, explore the same model catalogue plus the desktop-class OpenAI gpt-oss-20b.
No account. No server. Supported downloaded models can run without an internet connection.
LemoLM is a local-first playground for running, testing and comparing on-device language models, with detailed performance information for each generation.
FREE, NO ACCOUNT NEEDED
• Qwen3 1.7B — multilingual, with optional thinking mode
• SmolLM2 1.7B — compact, for older and lower-memory devices
• Apple Foundation Models — on supported devices
• Mock Engine — explore the interface without downloading anything
UNLOCK THE FULL MODEL PLAYGROUND
One purchase, no subscription. Adds access to twelve additional supported GGUF models:
• OpenAI gpt-oss-20b — desktop-class local reasoning for suitable high-memory Apple silicon Macs
• IBM Granite 4.2 3B — compact chat, coding and structured answers
• LFM2-350M ENJP-MT — fast English–Japanese and Japanese–English translation
• Mistral 7B Instruct v0.3 — general purpose and multilingual
• DeepSeek R1 Distill Qwen 7B — reasoning for high-memory devices
• DeepSeek R1 Distill Qwen 1.5B — compact reasoning
• Qwen3 4B — larger multilingual model with thinking mode
• NVIDIA Nemotron 3 Nano 4B — reasoning and coding
• Phi-4 mini Instruct — instruction following and reasoning
• Ministral 3 3B — chat, coding and instruction following
• SmolLM3 3B — compact general-purpose model
• TinySwallow 1.5B — Japanese-focused model from Sakana AI
SEE WHAT’S ACTUALLY HAPPENING
Most on-device AI apps hide the details. LemoLM shows them.
• Performance Snapshot — measured loading and generation timing, token and speed estimates, device and operating-system details, memory readings, thermal state, battery and power context
• Load Diagnostics — device memory, runtime mode and model file status when a model will not load
• Full model details — size, runtime, quantisation, licence and source
CONTROL THE OUTPUT
• Precise, Balanced and Creative presets
• Custom Temperature, Top-p, Top-k and Seed controls
• Adjustable reasoning effort on supported models
• Persistent settings with one-tap reset
COMPARE AND KEEP
• Rate models with your own stars and notes
• Export a Performance Snapshot and conversation together, or export either separately, as Markdown
• Learn about local AI concepts in the LLM Glossary
PRIVATE BY DESIGN
Conversations are stored on your device, and supported local models process prompts and responses on-device. Downloading a model requires internet access and may connect to third-party hosts such as Hugging Face. Once installed, supported downloadable models run locally.
LemoLM collects no data.
BEFORE YOU DOWNLOAD
Performance depends on your device, available memory, operating system version, thermal state and model size.
Compact models suit older or lower-me
On a suitable Apple silicon Mac, explore the same model catalogue plus the desktop-class OpenAI gpt-oss-20b.
No account. No server. Supported downloaded models can run without an internet connection.
LemoLM is a local-first playground for running, testing and comparing on-device language models, with detailed performance information for each generation.
FREE, NO ACCOUNT NEEDED
• Qwen3 1.7B — multilingual, with optional thinking mode
• SmolLM2 1.7B — compact, for older and lower-memory devices
• Apple Foundation Models — on supported devices
• Mock Engine — explore the interface without downloading anything
UNLOCK THE FULL MODEL PLAYGROUND
One purchase, no subscription. Adds access to twelve additional supported GGUF models:
• OpenAI gpt-oss-20b — desktop-class local reasoning for suitable high-memory Apple silicon Macs
• IBM Granite 4.2 3B — compact chat, coding and structured answers
• LFM2-350M ENJP-MT — fast English–Japanese and Japanese–English translation
• Mistral 7B Instruct v0.3 — general purpose and multilingual
• DeepSeek R1 Distill Qwen 7B — reasoning for high-memory devices
• DeepSeek R1 Distill Qwen 1.5B — compact reasoning
• Qwen3 4B — larger multilingual model with thinking mode
• NVIDIA Nemotron 3 Nano 4B — reasoning and coding
• Phi-4 mini Instruct — instruction following and reasoning
• Ministral 3 3B — chat, coding and instruction following
• SmolLM3 3B — compact general-purpose model
• TinySwallow 1.5B — Japanese-focused model from Sakana AI
SEE WHAT’S ACTUALLY HAPPENING
Most on-device AI apps hide the details. LemoLM shows them.
• Performance Snapshot — measured loading and generation timing, token and speed estimates, device and operating-system details, memory readings, thermal state, battery and power context
• Load Diagnostics — device memory, runtime mode and model file status when a model will not load
• Full model details — size, runtime, quantisation, licence and source
CONTROL THE OUTPUT
• Precise, Balanced and Creative presets
• Custom Temperature, Top-p, Top-k and Seed controls
• Adjustable reasoning effort on supported models
• Persistent settings with one-tap reset
COMPARE AND KEEP
• Rate models with your own stars and notes
• Export a Performance Snapshot and conversation together, or export either separately, as Markdown
• Learn about local AI concepts in the LLM Glossary
PRIVATE BY DESIGN
Conversations are stored on your device, and supported local models process prompts and responses on-device. Downloading a model requires internet access and may connect to third-party hosts such as Hugging Face. Once installed, supported downloadable models run locally.
LemoLM collects no data.
BEFORE YOU DOWNLOAD
Performance depends on your device, available memory, operating system version, thermal state and model size.
Compact models suit older or lower-me
Latest updates
1.7.1
Version 1.7.1 improves reliability and polish following the major 1.7 update.
• Improved reliability for consecutive LFM2 English–Japanese translations
• Improved GGUF generation and cancellation handling
• Expanded GPT-OSS response limits for more complete answers
• Improved chat presentation and conversation controls on Mac
• General stability and interface improvements
1.7.0
Version 1.7.0 expands LemoLM with three new local models and richer performance diagnostics:
• OpenAI GPT-OSS 20B brings desktop-class local reasoning to suitable high-memory Apple silicon Macs
• IBM Granite 4.2 3B adds a compact model for chat, coding and structured answers
• LFM2-350M ENJP-MT provides fast, private English–Japanese and Japanese-English translation
• Performance Snapshot now includes device, memory, thermal, battery and power context alongside generation timing
• Choose between complete, performance-only and conversation-only Markdown exports
• New model badges and an automatic What’s New summary make additions easier to discover
This release also includes interface refinements and stability improvements across iPhone, iPad and Apple silicon Mac.