AIPhonesComputers

Reddit User Splits AI Workload Between MacBook Pro M4 and iPhone 17 Pro Max via GitHub Project

The project offloads AI token layers to the iPhone's GPU, delivering up to 46% faster agent sessions, according to howtogeek.com.

A Reddit user going by StayLameBro has released a project that splits AI model tasks between a MacBook Pro M4 and an iPhone 17 Pro Max, according to howtogeek.com. The code, setup instructions, and benchmarking scripts are available on GitHub.

The setup divides processing of AI tokens—specifically for the Qwen 3.8 27B model—so that the Mac handles the first 40 layers of each token batch while the iPhone's GPU processes the remaining layers.

Benchmarks show the combination is between 29% and 44% faster than the Mac alone when prefilling a 2,000-token file into an existing agent session. For a new 27,000-token agent session, gains range from 36% for the project's forked llama.cpp code to 46% versus the stock software.

Hardware requirements include a standard 10Gbps USB-C cable and at least an iPhone 15 Pro.

Quick answers

What is needed to run this?

A MacBook Pro M4, an iPhone 17 Pro Max, a standard 10Gbps USB-C cable, and at least an iPhone 15 Pro.

How much faster is it?

Between 29% and 44% faster than the Mac alone for a 2,000-token prefill, and between 36% and 46% faster for a 27,000-token agent session.

Where can I find the code?

The code, setup process, and benchmarking scripts are in a GitHub repository.

Source