Apple is super stingy about RAM capacity on their devices, which is a problem for GenAI
Also, the M1 NPU (the equivalent to Google's offloading) is not very fast in Apple's own Stable Diffusion port, and I'm not aware of any community LLM frameworks that even use the NPU. They all run Metal implementations on the M1's GPU.