Dual M4 Mac Setup with Exo and Qwen3.5-9B

Eric Walker

I’ve been experimenting with local LLM inference, and after trying various setups, I finally landed on a sweet combination that delivers great performance for my coding work. Here’s the setup that got me to 100% satisfaction.

The Hardware: Two M4 Macs

Machine 1: M4 MacBook with 24GB unified memory Machine 2: M4 Mac with 36GB unified memory
Total: 60GB unified memory across two devices

Yes, that’s right—two Macs. Initially I thought this was overkill, but here’s why it makes sense:

  • Thunderbolt connectivity: Both machines are connected via Thunderbolt cable for low-latency communication
  • Redundant GPUs: Each M4 has its own GPU, allowing parallel processing when needed
  • Separation of concerns: I run Exo on each device independently, which gives me flexibility in workloads

The Software: Exo + Qwen3.5-9B

Why Exo?

Exo is a distributed inference framework that allows running models across multiple devices. What sets it apart:

  • Zero-copy RPC: Uses object graph virtualization (OGV) for near-instant data transfer between machines
  • GPU utilization: Efficiently leverages both M4 GPUs without complicated setup
  • Simple installation: Just cargo install exo and you’re good to go
  • Language support: I was using the Python client, which makes scripting straightforward

Why Qwen3.5-9B?

I chose the 9-billion parameter model with a clever quantization trick:

mlx-community/Qwen3.5-9B-8bit

The 8-bit quantization is key here. Here’s why:

  1. Memory efficiency: The model fits comfortably in RAM on a single device
  2. Performance: 8-bit is barely any loss compared to FP16 for most use cases
  3. Speed: Running at 8-bit means faster token generation than lower quantizations
  4. Quality retention: For coding tasks, the difference between 8-bit and full precision is negligible

Performance Characteristics

Token Generation Speed

With the dual-M4 setup, I’m seeing impressive speeds:

  • Single device (24GB Mac): ~25-30 tokens/second
  • Single device (36GB Mac): ~35-40 tokens/second
  • Distributed across both: The individual devices handle their own workloads efficiently

For my use case (coding assistant), this means I’m getting near-instant responses to questions and code completions. The model never feels sluggish, even during complex reasoning tasks.

Context Window

Running Qwen3.5-9B with its native context window means I can:

  • Feed entire codebases for analysis
  • Maintain full conversation history without truncation
  • Process large context-heavy coding tasks smoothly

Installation and Configuration

Prerequisites

  • M4 Mac (verified setup)
  • Rust toolchain for Exo backend
  • Python 3.10+ for the client

Setup Steps

# Install Exo backend on both devices
cargo install exo

# Start each device's inference server
exo serve --model mlx-community/Qwen3.5-9B-8bit

# Configure client to use distributed setup
export EXO_DISTRIBUTED=true

Client Configuration

For Python clients, configuration is straightforward:

from exo import Client

client = Client(
    model="mlx-community/Qwen3.5-9B-8bit",
    distributed=True,
    devices=["device-1.local", "device-2.local"]
)

response = client.complete("Explain this code...")

Why This Setup Wins

1. Local Privacy First

Everything runs locally. No cloud API calls, no data leaving my machines. For coding work where I paste proprietary code or business logic, this is non-negotiable.

2. Cost-Effective Performance

Running Qwen3.5-9B costs nothing (it’s an open model, available on HuggingFace), and I’m using hardware I already own. Compare this to cloud APIs charging $0.02-0.05 per 1K tokens—this local setup is infinitely more cost-effective at scale.

3. Deterministic Response Quality

Local inference means I don’t have to worry about:

  • API rate limits or downtime
  • Variable model versions or configurations
  • Unexpected token costs for large projects

4. Development Workflow Integration

The latency is fast enough to feel instant, so I can use it as a true pair programmer. No longer do I have to “wait my turn” while the model generates responses—I’m in a natural conversational flow.

What I’m Using It For

  • Code completion and explanation: Real-time suggestions while coding
  • Bug detection: Feed error messages for rapid diagnosis
  • Documentation generation: Convert code comments to docstrings automatically
  • Code review: Ask for explanations of unfamiliar patterns
  • Learning: Get instant answers to “why does this API work this way?”

Future Enhancements

As I experiment more, here are potential improvements:

  • Model switching: Having the flexibility to swap between models based on task complexity
  • Batch inference: Running multiple queries in parallel for efficiency
  • Fine-tuning: Customizing the model on my coding patterns and preferences

Conclusion

The combination of two M4 Macs, Exo for distributed inference, and Qwen3.5-9B (8-bit quantized) creates a local LLM setup that’s both performant and practical for everyday coding tasks. It’s fast, private, cost-free, and integrates seamlessly into my development workflow.

If you’re looking for a local LLM setup that balances quality, speed, and cost, give this combination a try. You might be surprised by how capable even an “only” 9B parameter model can be when tuned for your use case.


Tags: #localLLM #Exo #Qwen3.5 #MLX #M4Mac #AI