Hacker News new | past | comments | ask | show | jobs | submit
Is there any relevant model that meets that "open source" definition?
No major language model I am aware of meets the polar extreme that I describe as unmistakably open source, because even those with transparent training data (like IBM Granite) generally do kot use exclusively training data that either they own and can control the license, are under an open license, or are public domain.

OTOH, to the extent that the original model trainers rely on training on certain data not requiring a license from the copyright holder, there is at least an argument that with an open source licenses for the weights and training and inference code, a transparent training corpus to which the original trainer has relied on no special permissions not granted to the general public to train on it, to the extent that the legal theory behind the original trainer believing that it is free to train on the data is correct, provides all of the essential features of open source.

At the same time, there are things portrayed as open weights where training data is undisclosed and the weights have a license which limits purpose of use and other aspects of use; the models are free-of-charge (for limited uses) but not meaningfully open.

It depends on if you count https://allenai.org/ as relevant?
loading story #49253911