Learning Local AI: Exciting, Overwhelming, Frustrating
Antonio G. Di Benedetto of The Verge begins testing local AI models using an M5 Ultra Mac Studio and Hermes Agent for data privacy.

Stock photo for illustration only, not from the actual event
- Antonio G. Di Benedetto starts testing local AI hardware processing to protect personal data privacy.
- He utilizes the open-source Hermes Agent desktop app on an M5 Ultra Mac Studio with 256GB of unified memory.
- The test begins with a massive Qwen 3.8 Flash Next model featuring 125 billion parameters and a 105GB size.
- Integration with a Telegram bot allows remote control of the local AI agent directly from a smartphone.
Using cloud-based artificial intelligence tools often comes with the nagging concern of sharing deeply personal data, prompting The Verge reviewer to take the plunge into running powerful AI models directly on local hardware. The goal is to explore what the technology is truly capable of without relying on external cloud servers, even if it requires investing in expensive computers or laptops with hefty RAM upgrades. The author openly admits to not being an AI expert, embarking on this journey alongside readers to determine whether the effort is truly worthwhile.
Local AI and autonomous agents are experiencing a massive surge in popularity, heavily pushed by Apple's marketing for new Mac desktops and upcoming Windows RTX Spark machines boasting up to 128GB of RAM. The appealing concept of having a personal digital assistant living entirely inside a box on the desk—one that answers solely to its owner rather than corporate overlords at OpenAI, Google, Microsoft, or Anthropic—offers a stark contrast to trusting subscription-based cloud chatbots.

Stock photo for illustration only, not from the actual event
As part of ongoing evaluations using the M5 Ultra Mac Studio alongside various other Mac and Windows systems, the author installed several local large language models (LLMs) to begin experimenting. The process kicked off with the open-source Hermes Agent, a self-hosted desktop application compatible with macOS, Windows, and Linux that remains completely free when paired with local LLMs. Given that the test machine boasts an impressive 256GB of unified memory, it possesses the capability to run practically any available model.
Running large language models locally is emerging as a critical approach for users prioritizing strict data privacy, ensuring sensitive inputs never leave their own hardware infrastructure, despite the steeper hardware costs and setup complexity.
This setup process quickly highlights how overwhelming the ecosystem can be, given the countless models available with varying specialized use cases. Generally, models featuring billions of parameters offer broader capabilities than smaller alternatives. Since local execution eliminates token fees, the author decided to start on a large scale using the straightforward onboarding interface provided by Hermes.
The initial choice landed on the Qwen 3.8 Flash Next model, a staggering 125-billion parameter architecture requiring approximately 105GB of storage space. Future testing plans involve experimenting with much smaller Qwen variants across alternative devices, including an M6 Mac Mini, an M5 MacBook Air, an Asus TUF Gaming A14 powered by an AMD Strix Halo processor, and eventually an RTX Spark system.
Getting Hermes and Qwen operational took minimal time, even extending control to a mobile device via a Telegram bot. However, this progress immediately led straight to the most familiar nemesis encountered when interacting with artificial intelligence systems: the empty text box.
Source: The Verge
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment