i have been following Chinese models for about two years now because they are open-weight and qwen is fun to run on my kubernetes cluster, the news about this Apache licensed model complete with a recipe to make it again with potentially different ingredients is making we want to abandon them for something more ideologically sound and way more interesting

this new model, couldn’t universities rebuild it with different contexts and study it in ways you can’t reproduce in other models? like, is k2 horizons a good scientific foundation on which to study machine learning?

or am i just lacking way too much context and falling for hype?

  • joseki@lemmy.zip
    link
    fedilink
    English
    arrow-up
    12
    ·
    16 hours ago

    K2 wasn’t the first.

    China had CPM-1 in 2020 that was a 2.6B model.

    GPT-Neo (Mar 2021) EleutherAI replicates early GPT

    PanGu-α (Apr 2021) Huawei’s 200B model

    WuDao 2.0 (May 2021) BAAI’s massive 1.75T sparse MoE model

    GPT-J (Jun 2021) EleutherAI release 6B

    Meta released OPT in 2022, BLOOM was in 2022, GLM 130B by Tsinghua/Zhipu was in 2022 etc

    K2 horizon is very recent, nowhere near the first. Different labs have different computational innovations worth studying. DeepSeek is crazy efficient and doing very novel things. Mistral in France is making very compact local friendly models that are fun and easy to fine tune and merge on consumer hardware.

    • hendrik@palaver.p3x.de
      link
      fedilink
      English
      arrow-up
      4
      ·
      8 hours ago

      Just to add a reference time frame: late 2022 was when ChatGPT got released. I think all these models predate the general population even being aware of AI.

      • joseki@lemmy.zip
        link
        fedilink
        English
        arrow-up
        3
        ·
        7 hours ago

        Right and the predecessor to ChatGPT was trained on links from reddit posts pretty wild. I remember having fun making it come up with nonsense prose on my GPU