A user has outlined their personal local LLM server configuration, which operates on an M4 Pro Mac mini equipped with 48 GB of RAM. This setup manages diverse tasks, ranging from an agent backend to immediate chat responses on a mobile phone. The entire system can be configured in approximately 30 minutes.
The core of the setup includes Qwen3.6-35B-A3B-OptiQ-4bit for complex reasoning tasks and Gemma-4-E4B-it-OptiQ-4bit for simpler chats and formatting. oMLX functions as the inference server, while Tailscale creates a secure network connecting the Mac mini, iPhone, and MacBook. Applications like Hermes, Apollo, Pi, and Raycast AI are integrated for various uses, including agent workflows, quick chats, and coding assistance.
The primary reasons for opting for a local setup include avoiding the variable pricing, usage limits, and unannounced model changes associated with cloud APIs. The user previously incurred significant monthly costs and experienced inconsistent model performance from cloud services. Data privacy is another critical concern, as sending sensitive information to third-party APIs introduces operational security risks and loss of control over data usage.
AI sovereignty is cited as a factor, highlighting the risk of government restrictions on cloud models that could disrupt workflows. Local compute provides independence from such external controls. Additional practical benefits include predictable costs, as hardware purchase and electricity are fixed expenses, making subsequent inferences free. Reduced latency due to the absence of network roundtrips and the ability to operate offline are also significant advantages for daily tasks and background agent workflows.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
A user detailed their personal local LLM setup running on an M4 Pro Mac mini with 48 GB of RAM, utilizing models like Qwen3.6-35B and Gemma-4-E4B with oMLX as the inference server. The setup supports various applications from agent backends to quick chat queries on mobile devices. The user cites cost predictability, data privacy, AI sovereignty, lower latency, and offline capability as primary motivations for running LLMs locally instead of relying on cloud APIs.