improve Your AI Setup with Multiple Machines
30 September 2026
Today, I was working in my studio in Munich, and I realised something important about our AI setup. It's a common issue that founders face when they start using local AI: at some point, one machine just isn't enough. This happens when you hit memory limits or need more concurrency for multiple agents running simultaneously.
When you reach this point, it's time to think about adding a second machine to your fleet. But before diving in, understand the practicalities of how these machines interact and what role each plays. For me, having two machines, one large and one smaller, has been essential.
The larger machine in my setup is a 64GB Mac Studio that sits in my office. This is the brain of our AI network. It runs the largest model I use, which is necessary for tasks like writing detailed business reports or summarising complex meetings. The Studio holds all the canonical business data and is always on to ensure continuous operation without interruptions.
Next to the Studio, there's a 16GB Mac Mini. This machine handles smaller, faster models that don't require as much memory. It runs background jobs like email classification, daily summaries, and routine drafting tasks. These processes are less critical in terms of immediate response time, so they can run on a smaller machine without affecting the performance of more demanding tasks.
My laptop, a MacBook Pro, acts differently. When it's open, it connects to either the Studio or the Mini for assistance. However, when closed, it doesn't disrupt anything because it's not a core part of the network. It's a client rather than an equal peer in the setup.
This architecture, having one brain (the Studio), one helper (the Mini), and one client (the laptop), has proven effective. Each machine has a specific role that aligns with its capabilities, ensuring that resources are used efficiently without overcomplicating things.
The practical question then becomes: how do you split the workload between these two machines? A straightforward approach is to base it on model size. Larger models run on the more powerful Studio, while smaller models use the Mini. For example, summarising a meeting can be handled by the Mini's smaller model, whereas writing an investor update might require the larger model on the Studio.
Another split involves task type. Long-running background jobs that don't need immediate attention run on the Mini, leaving the Studio free to handle interactive tasks. This keeps the chat functions fast and responsive without delays caused by competing processes.
Avoid over-engineering your setup initially. Start with a simple rule for splitting tasks between the two machines. As you gain experience, adjust the rules based on what works best in practice. In my experience, the complexity of managing more than three machines starts to outweigh the benefits. A third machine might add capacity but also introduces additional maintenance requirements.
There is one specific scenario where adding a fourth machine makes sense: when multiple people need simultaneous access to AI tools within an actual team setup. In such cases, a dedicated inference server that all members can connect to becomes necessary.
Understanding these principles helps in scaling your local AI network without making it overly complex. Remember, the goal is efficiency and effectiveness rather than sheer numbers of machines. By keeping things simple yet robust, you can use local AI to its fullest potential for your business needs.
So, next time you feel like one machine isn't cutting it anymore, consider adding a second but remember to define clear roles and rules. This approach ensures that your setup remains manageable and powerful, allowing you to focus on the bigger picture of growing your venture.
This piece was written by my AI editorial team: Sven scouted the topic, Ines gathered and verified sources, Linnea drafted the body, Vera fact checked every claim against the cited URLs, Bea edited for my voice, and Sora generated the hero image. All on a Mac in my Munich studio, no cloud. I read every piece before it goes live during the launch window. If something is wrong, write to me.
