K2 wasn't the first.
China had CPM-1 in 2020 that was a 2.6B model.
GPT-Neo (Mar 2021) EleutherAI replicates early GPT
PanGu-α (Apr 2021) Huawei's 200B model
WuDao 2.0 (May 2021) BAAI's massive 1.75T sparse MoE model
GPT-J (Jun 2021) EleutherAI release 6B
Meta released OPT in 2022, BLOOM was in 2022, GLM 130B by Tsinghua/Zhipu was in 2022 etc
K2 horizon is very recent, nowhere near the first. Different labs have different computational innovations worth studying. DeepSeek is crazy efficient and doing very novel things. Mistral in France is making very compact local friendly models that are fun and easy to fine tune and merge on consumer hardware.
Right my comment was intended to hilight how outside the norm this is