this post was submitted on 08 Sep 2026
11 points (92.3% liked)

Selfhosted

61978 readers
1222 users here now

A place to share alternatives to popular online services that can be self-hosted without giving up privacy or locking you into a service you don't control.

Rules:

Detailed Rules Post

  1. Be civil.

  2. No spam.

  3. Posts are to be related to self-hosting.

  4. Don't duplicate the full text of your blog or readme if you're providing a link.

  5. Submission headline should match the article title.

  6. No trolling.

  7. Promotion posts require active participation, with an account that is at least 30 days old. F/LOSS without a paywall has exceptions, with requirements. See the rules link for details. Tags [CBH] or [AIP] are required, see the links in Rule 8 for details.

  8. AI-related discussions and AI-involved promotional posts have additional requirements for tagging, as noted in Rule 7 and the AI & Promotional Post Expanded Rules post, and find example disclosures here.

Resources:

Any issues on the community? Report it using the report flag.

Questions? DM the mods!

founded 3 years ago
MODERATORS
 

Hi! So I'm considering...maybe having an NPU or something similar to be hooked to my proxmox server, which runs in a mini PC. It's a EliteDesk 800 micro form factor. It has a Core i5 8500 CPU, which at the moment of purchase was good enough for live encoding HEVC video on Jellyfin...that was my main concern back then. But I'd like to consider the possibility of hooking maybe some docker instances or other containers to some local-only AI acceleration. Is there any NPU or cheap GPU I could hook on USB to this proxmox server to run? Has it been done before?

Thanks!

you are viewing a single comment's thread
view the rest of the comments
[–] Sims@lemmy.ml 4 points 6 hours ago

USB is not good as its a huge bottleneck, and most external accelerators are (was?) passive without any ram onboard. Google Coral was/is 'passive' in that it resends all data all the time over the interface - Npu<->system ram.

I were going to recommend these: https://shop.geniatech.com/product/m2-ai-inference-acceleration-module/

40tops, ARM + Npu + 16gb ddr4 - a whole little Inference computer on a NVME interface. Kinara (Ara240) is a homegrown Chinese chip. While they are usually selling b2b, you can ask anyway. YMMW atmo. Also note that some NPU's are less efficient at llm's vs vision.

..but I see that they also rose almost 4* in price since I asked for, and were offered the price of 179$ ~7M ago, which is already a long time in this space. Not sure how the current ~650$ stacks up to the rest of the offers out there at the moment, but these small active AI systems on NVME are a great way of upgrading a piss-old server, enhance a new cheap 4-8port NVME mini-NAS or similar, and there's no clear bottleneck in the interface.

Look for something like this instead of Nvidia Coral and other 'passive' sticks, that are all - imho - overpriced/underperforming.