this post was submitted on 24 Jun 2026
76 points (79.7% liked)

Selfhosted

60093 readers
858 users here now

A place to share alternatives to popular online services that can be self-hosted without giving up privacy or locking you into a service you don't control.

Rules:

  1. Be civil: we're here to support and learn from one another. Insults won't be tolerated. Flame wars are frowned upon.

  2. No spam.

  3. Posts here are to be centered around self-hosting. Please ensure it is clear in your post how it relates to self-hosting.

  4. Don't duplicate the full text of your blog or git here. Just post the link for folks to click.

  5. Submission headline should match the article title.

  6. No trolling.

  7. Promotion posts require your active participation in selfhosting or related communities, or the post will be removed. No more than 10% of your posts or comments may be self-promotional, or your post will be removed. F/LOSS Exception: If your post is about a project that is completely open source & can be self-hosted in full without payment, your post is exempt from this rule as long as you continue to engage in comments.

Resources:

Any issues on the community? Report it using the report flag.

Questions? DM the mods!

founded 3 years ago
MODERATORS
 

Do you host your own ML / AI / LLM? What do you use, and what do you use it for?

you are viewing a single comment's thread
view the rest of the comments
[–] brucethemoose@lemmy.world 1 points 10 hours ago (1 children)

How much CPU RAM do you have?

[–] atzanteol@sh.itjust.works 1 points 10 hours ago (1 children)

64G. But CPU inference is painfully slow.

[–] brucethemoose@lemmy.world 6 points 9 hours ago* (last edited 9 hours ago) (1 children)

Not anymore. Not with hybrid offloading, where the GPU handles dense tensors and the CPU only runs the sparse MoEs. I'm running a 300B model on a single 3090, and its faster than I can read.

You just need to use the right framework, and the right model.

I'd suggest trying ik_llama.cpp and a MoE like one of these: https://huggingface.co/models?other=ik_llama.cpp&sort=modified&search=35B

And speculative decoding like DFlash or MTP (which you can also get specific models for).

EDIT: Wrong link.

[–] atzanteol@sh.itjust.works 1 points 8 hours ago (1 children)

I'll check that out - speed isn't my biggest issue so much as coding performance... The qwen 3.5 model I was using can write code, but it's... Meh? Like sometimes it doesn't even compile.

I did try tweaking llama.cpp to do some cpu offloading and it does seem to allow for much larger contexts at a modest performance loss. I'll check out larger models.

[–] brucethemoose@lemmy.world 1 points 7 hours ago* (last edited 7 hours ago)

CPU offloading is too slow unless you use a hybrid MoE model, with the --n-cpu-moe parameter, specifically.

This only offloads "sparse" parts of the model to the CPU, which take up a lot of RAM but are very compute-lite to run. In practice, thats most of the size of modern MoE LLMs.