USB is not good as its a huge bottleneck, and most external accelerators are (was?) passive without any ram onboard. Google Coral was/is 'passive' in that it resends all data all the time over the interface - Npu<->system ram.
I were going to recommend these: https://shop.geniatech.com/product/m2-ai-inference-acceleration-module/
40tops, ARM + Npu + 16gb ddr4 - a whole little Inference computer on a NVME interface. Kinara (Ara240) is a homegrown Chinese chip. While they are usually selling b2b, you can ask anyway. YMMW atmo. Also note that some NPU's are less efficient at llm's vs vision.
..but I see that they also rose almost 4* in price since I asked for, and were offered the price of 179$ ~7M ago, which is already a long time in this space. Not sure how the current ~650$ stacks up to the rest of the offers out there at the moment, but these small active AI systems on NVME are a great way of upgrading a piss-old server, enhance a new cheap 4-8port NVME mini-NAS or similar, and there's no clear bottleneck in the interface.
Look for something like this instead of Nvidia Coral and other 'passive' sticks, that are all - imho - overpriced/underperforming.