There would a lot of problems with the feasibility and enforceability of such a project, but putting that aside…

    • Skullgrid@lemmy.world
      link
      fedilink
      arrow-up
      9
      arrow-down
      1
      ·
      2 months ago

      the problems with AI can be easily mitigated, they just aren’t for various reasons.

      You can have an AI trained on public domain content, running on electricity from renewable sources and locally hosted. What’s the problem then?

      • schipelblorp@sh.itjust.works
        link
        fedilink
        arrow-up
        4
        arrow-down
        4
        ·
        2 months ago

        “Various reasons” being the intractability and absolute dominance of capitalism. People are in a race to dominate the world, as such they can’t be bothered with writing public domain software, or waiting for renewable energy to come online. They can’t even wait for hardware to be built before buying it!

        I mean, it’s kind of like asking a herd of zebras, “would lions be so bad if they just ate grass instead?” Like, sure, but the reality is that lions will always eat meat just as inevitably as AI will always be horrible under capitalism.

  • SnailMagnitude@mander.xyz
    link
    fedilink
    English
    arrow-up
    11
    arrow-down
    1
    ·
    2 months ago

    yeah, it would be a large use of resources for not much gain

    I think what the current LLM clusterfuck is showing is that our little world or copy left and copy right we duct taped together since the printing press is all a bit silly.

  • floquant@lemmy.dbzer0.com
    link
    fedilink
    arrow-up
    5
    ·
    2 months ago

    I keep thinking that we would be in a much better place if datasets were curated and models trained by respectable institutions like universities and foundations, I would “hate AI” a lot less and just use the LLMs that themselves have done nothing wrong. (Not anthropomorphizing, meaning “they’re just a bunch of quite cool math”)

  • communism@lemmy.ml
    link
    fedilink
    arrow-up
    5
    arrow-down
    1
    ·
    2 months ago

    I think the “copyright” shit is the biggest non-issue with LLMs. Abolish intellectual property. If you put a piece of software out there, you shouldn’t be able to stop anyone else from using it for the public good.

    The issues I take with what’s being called “AI” right now is the volume of slop people have to drudge through. It particularly makes it hard to look for new software as there’s so many vibe-coded projects getting published then abandoned. It’s the aggressive and obnoxious marketing campaign that insists I pay for your shitty SaaS LLM, that I don’t get to keep or run on my own machine, that I’m supposed to use for all sorts of tasks where I don’t want a non-deterministic probabilistic machine (people use LLMs to do number-crunching?? you’re using a machine designed to do math, interfacing with the one part of it that can’t do math!). And it’s the ridiculous computing power people are using to do activities that do not require anywhere near that amount of power if they just do things the normal way. I could not care less about “copyright”. Nobody should be able to own a codebase anyway.

    • BlameThePeacock@lemmy.ca
      link
      fedilink
      English
      arrow-up
      3
      ·
      2 months ago

      Agreed, the concept of copyright itself is problematic. The flip on this one boggles me. People are constantly arguing that copyright is bad every time Disney or some company use it to make money. Now they’re arguing that copyright is good and LLMs should respect it.

  • vala@lemmy.dbzer0.com
    link
    fedilink
    arrow-up
    3
    ·
    2 months ago

    Something like this would be interesting because it would make the other models seem even less ethical.

  • litchralee@sh.itjust.works
    link
    fedilink
    English
    arrow-up
    2
    ·
    edit-2
    2 months ago

    Given that FOSS licenses are premised on copyright, yes, the same ails would still exist: 1) AI washing of licenses (including transforming one license into another), and 2) the vagueness of whether LLM outputs can be copyright, which threatens the validity of a FOSS license upon that output.

    The first ail can be seen even without LLMs: the BSD variants have gone through great pains to remove GPL-licensed code from their base repositories. This basically involves reimplementing utilities and functionality from scratch, using only the ideas that are in common with the equivalent GPL code, but never copying that code directly. This is properly considered a reimplementation, which can then be licensed permissively (eg MIT license).

    If an LLM were to train on GPL code but the output were licensed with MIT, then that could be a GPL violation because GPL mandates that remixes continue to keep the GPL license.

    Maybe you could avoid this fate by limiting the LLM to only train on permissively licensed code. So that it would be permissive licenses going in, and permissive licenses coming out. No GPL problems here. But that brings us to ail #2.

    Some jurisdictions have rules against granting copyright for computer-generated works, in the same vein as works generated by non-humans (eg a macaque). If this LLM fell into this situation, then the output is not copyrightable. And if there is no copyright, a license like MIT or GPL simply cannot apply, because its terms couldn’t be enforced.

    Well, to be clear, the copyright parts of those licenses would be unenforceable. Some parts of the license may still be enforced under a contracts claim. But in any case, the things we refer to as “FOSS licenses” cannot attach to uncopyrightable works (with the possible exception of the CC0 license, which is essentially the absence of any license whatsoever).

    EDIT: you did say “consenting projects” and consent is key. If such consent came in the form of a license grant, then yes, that would be enthusiastic consent for the LLM to generate output, which solves ail #1. But for most multi-person projects, getting consent from everyone is difficult or impossible. The Linux kernel is one such example, having so many contributors that some of them are already dead. Death means they cannot consent, but their copyright grant lives on. And so practically speaking, obtaining enthusiastic consent for whole projects is a challenge, which drastically limits the prospects for such an LLM from the very beginning.

  • hperrin@lemmy.ca
    link
    fedilink
    English
    arrow-up
    1
    ·
    2 months ago

    It would still be problematic, yes. You can’t copyright the output of an LLM, so you can’t actually enforce any license on it.