Some advice I got from another HN Mac user was to run local models in energy saver mode. You'll get slightly reduced tokens, but the laptop won't overheat and the fans won't go wild.
And if you're running it on a dGPU, power limit it, because you lose very little in terms of token generation performance, since it's memory-bound.