this post was submitted on 20 Sep 2026
34 points (82.7% liked)

Selfhosted

62234 readers
766 users here now

A place to share alternatives to popular online services that can be self-hosted without giving up privacy or locking you into a service you don't control.

Rules:

Detailed Rules Post

  1. Be civil.

  2. No spam.

  3. Posts are to be related to self-hosting.

  4. Don't duplicate the full text of your blog or readme if you're providing a link.

  5. Submission headline should match the article title.

  6. No trolling.

  7. Promotion posts require active participation, with an account that is at least 30 days old. F/LOSS without a paywall has exceptions, with requirements. See the rules link for details. Tags [CBH] or [AIP] are required, see the links in Rule 8 for details.

  8. AI-related discussions and AI-involved promotional posts have additional requirements for tagging, as noted in Rule 7 and the AI & Promotional Post Expanded Rules post, and find example disclosures here.

Resources:

Any issues on the community? Report it using the report flag.

Questions? DM the mods!

founded 3 years ago
MODERATORS
 

i have been following Chinese models for about two years now because they are open-weight and qwen is fun to run on my kubernetes cluster, the news about this Apache licensed model complete with a recipe to make it again with potentially different ingredients is making we want to abandon them for something more ideologically sound and way more interesting

this new model, couldn't universities rebuild it with different contexts and study it in ways you can't reproduce in other models? like, is k2 horizons a good scientific foundation on which to study machine learning?

or am i just lacking way too much context and falling for hype?

you are viewing a single comment's thread
view the rest of the comments
[–] Natanox@discuss.tchncs.de 2 points 1 hour ago

I mean, they "plan" to publish that info, right?

I'd bet they used at least CommonCrawl (with it being the majority of data), arXiv, Wikipedia and the set containing all of Github. Probably not a lot of distillation.

CommonCrawl is one of the reasons small websites and social instances get DDoS'ed by rules-ignoring AI crawlers. Wikipedia data basically always gets used without paying them. Github… well, it's a prime example of how they broke millions of licenses.

There's also other stuff commonly used. Don't get me started on the training sets for image generation. It's beyond disgusting (and I'm not even talking just about theft and cultural destruction at this point).

This stuff is a bottomless pit, and if those (F)OSS models really aim to be up at the top they'll have to break every possible moral, ethical and legal rule just like everyone else. Even more so if they omit closed-source training sets.