I’ve been experimenting with local LLM inference, and after trying various setups, I finally landed on a sweet combination that delivers great performance for my coding work. Here’s the setup that got me to 100% satisfaction.
The Hardware: Two M4 Macs
Machine 1: M4 MacBook with 24GB unified memory
Machine 2: M4 Mac with 36GB unified memory
Total: 60GB unified memory across two devices
Yes, that’s right—two Macs. Initially I thought this was overkill, but here’s why it makes sense:
- Thunderbolt connectivity: Both machines are connected via Thunderbolt cable for low-latency communication
- Redundant GPUs: Each M4 has its own GPU, allowing parallel processing when needed
- Separation of concerns: I run Exo on each device independently, which gives me flexibility in workloads
The Software: Exo + Qwen3.5-9B
Why Exo?
Exo is a distributed inference framework that allows running models across multiple devices. What sets it apart:
- Zero-copy RPC: Uses object graph virtualization (OGV) for near-instant data transfer between machines
- GPU utilization: Efficiently leverages both M4 GPUs without complicated setup
- Simple installation: Just
cargo install exoand you’re good to go - Language support: I was using the Python client, which makes scripting straightforward
Why Qwen3.5-9B?
I chose the 9-billion parameter model with a clever quantization trick:
mlx-community/Qwen3.5-9B-8bit
The 8-bit quantization is key here. Here’s why:
- Memory efficiency: The model fits comfortably in RAM on a single device
- Performance: 8-bit is barely any loss compared to FP16 for most use cases
- Speed: Running at 8-bit means faster token generation than lower quantizations
- Quality retention: For coding tasks, the difference between 8-bit and full precision is negligible
Performance Characteristics
Token Generation Speed
With the dual-M4 setup, I’m seeing impressive speeds:
- Single device (24GB Mac): ~25-30 tokens/second
- Single device (36GB Mac): ~35-40 tokens/second
- Distributed across both: The individual devices handle their own workloads efficiently
For my use case (coding assistant), this means I’m getting near-instant responses to questions and code completions. The model never feels sluggish, even during complex reasoning tasks.
Context Window
Running Qwen3.5-9B with its native context window means I can:
- Feed entire codebases for analysis
- Maintain full conversation history without truncation
- Process large context-heavy coding tasks smoothly
Installation and Configuration
Prerequisites
- M4 Mac (verified setup)
- Rust toolchain for Exo backend
- Python 3.10+ for the client
Setup Steps
# Install Exo backend on both devices
cargo install exo
# Start each device's inference server
exo serve --model mlx-community/Qwen3.5-9B-8bit
# Configure client to use distributed setup
export EXO_DISTRIBUTED=true
Client Configuration
For Python clients, configuration is straightforward:
from exo import Client
client = Client(
model="mlx-community/Qwen3.5-9B-8bit",
distributed=True,
devices=["device-1.local", "device-2.local"]
)
response = client.complete("Explain this code...")
Why This Setup Wins
1. Local Privacy First
Everything runs locally. No cloud API calls, no data leaving my machines. For coding work where I paste proprietary code or business logic, this is non-negotiable.
2. Cost-Effective Performance
Running Qwen3.5-9B costs nothing (it’s an open model, available on HuggingFace), and I’m using hardware I already own. Compare this to cloud APIs charging $0.02-0.05 per 1K tokens—this local setup is infinitely more cost-effective at scale.
3. Deterministic Response Quality
Local inference means I don’t have to worry about:
- API rate limits or downtime
- Variable model versions or configurations
- Unexpected token costs for large projects
4. Development Workflow Integration
The latency is fast enough to feel instant, so I can use it as a true pair programmer. No longer do I have to “wait my turn” while the model generates responses—I’m in a natural conversational flow.
What I’m Using It For
- Code completion and explanation: Real-time suggestions while coding
- Bug detection: Feed error messages for rapid diagnosis
- Documentation generation: Convert code comments to docstrings automatically
- Code review: Ask for explanations of unfamiliar patterns
- Learning: Get instant answers to “why does this API work this way?”
Future Enhancements
As I experiment more, here are potential improvements:
- Model switching: Having the flexibility to swap between models based on task complexity
- Batch inference: Running multiple queries in parallel for efficiency
- Fine-tuning: Customizing the model on my coding patterns and preferences
Conclusion
The combination of two M4 Macs, Exo for distributed inference, and Qwen3.5-9B (8-bit quantized) creates a local LLM setup that’s both performant and practical for everyday coding tasks. It’s fast, private, cost-free, and integrates seamlessly into my development workflow.
If you’re looking for a local LLM setup that balances quality, speed, and cost, give this combination a try. You might be surprised by how capable even an “only” 9B parameter model can be when tuned for your use case.
Tags: #localLLM #Exo #Qwen3.5 #MLX #M4Mac #AI