Selfhosted
A place to share alternatives to popular online services that can be self-hosted without giving up privacy or locking you into a service you don't control.
Rules:
-
Be civil.
-
No spam.
-
Posts are to be related to self-hosting.
-
Don't duplicate the full text of your blog or readme if you're providing a link.
-
Submission headline should match the article title.
-
No trolling.
-
Promotion posts require active participation, with an account that is at least 30 days old. F/LOSS without a paywall has exceptions, with requirements. See the rules link for details. Tags [CBH] or [AIP] are required, see the links in Rule 8 for details.
-
AI-related discussions and AI-involved promotional posts have additional requirements for tagging, as noted in Rule 7 and the AI & Promotional Post Expanded Rules post, and find example disclosures here.
Resources:
- selfh.st Newsletter and index of selfhosted software and apps
- awesome-selfhosted software
- awesome-sysadmin resources
- Self-Hosted Podcast from Jupiter Broadcasting
Any issues on the community? Report it using the report flag.
Questions? DM the mods!
view the rest of the comments
Do you want to run TensorFlow Lite / LiteRT models? PyTorch Mobile? TensorRT? onnx? YOLO? vLLM? Something else? The recommendations will vary based on your use case.
Google Coral was decent for TensorFlow Lite, but it's EOL (end of life) now. I've got the dual TPU Mini PCIe version in my home server, via a PCIe adapter board. I use it for object detection with Blue Iris + CodeProject AI and it works pretty well for that use case.
Hailo-8 is supposed to be like a more powerful version of the Coral, but I don't have experience with it. It supports a bunch of frameworks: TensorFlow, TensorFlow Lite, Keras, PyTorch and ONNX. I'd be interested in hearing other people's thoughts on it.
I don't know if any of these work over USB though. They're usually internal devices. Google marketed the Coral USB as being for development and testing only, pointing people to the M.2 and PCIe versions for production usage.
As for something totally different... There's the Nvidia Jetson single board computer which supports TensorRT, but I don't have experience with it either. I also think it's a bit older too. You could also consider getting a newer mini PC with a AMD Ryzen AI processor in it, or an Nvidia DGX Spark.
Google's latest TPUs are only available in their cloud - they're not selling the hardware to end users any more.
Thanks! I was considering using it perhaps to have a diffusion model? Or running ollama or similar without a RAM hit on my laptop or phone. Also whisper comes to mind, for Bazarr or other tools to use.
Any NPU/TPU you can buy is going to be essentially useless for either image diffusion or LLMs. The onboard RAM is both far too small and far too slow (LLM text generation speed relies on RAM speed first and foremost, and both LLMs and image models tend to be, you know, big), and USB isn't nearly fast enough to help with that, not to mention that software support is pretty much nonexistent. You'd be better off upgrading the GPU to a 3060ti 12gb or something.
P.S. A word of advice, consider using something other than Ollama. Llama.cpp in router mode or llama-swap support pretty much all of the functionality that Ollama does without being crap. Ik_llama.cpp is also nice if you have a CPU/Nvidia setup.
If you wanna make the most out of what you've got now, the LFM2.5 series of LLMs are quite good for the small size and fast inference speeds with sizes ranging from 0.2 billion to 8 billion parameters, though their low parameter count means that you'll probably wanna hook them up to some sort of web search or similar since they won't have a ton of general knowledge.
If you have at least 32GB of RAM, Qwen3.6 35B is quite a good general-purpose model that runs faster than its parameter count would suggest.
Aren't diffusion models and LLMs (ollama) too big for an external NPU? As far as I know something like a Coral runs specific models only. And it's limited to the 1 or maybe 2GB of memory on it. It'd do tasks like voice recognition, or image classification. But not generate images or text.
If you want to run arbitrary AI models and generative AI, I think you should be looking for a graphics card?!
Agreed. OP should probably upgrade to a bigger case and discrete graphics card.