Updated
Updated · PBS NewsHour · Jul 25
Researcher Distinguishes 3 AI Model Types as Open-Source Debate Turns on Code, Weights and Data
Updated
Updated · PBS NewsHour · Jul 25

Researcher Distinguishes 3 AI Model Types as Open-Source Debate Turns on Code, Weights and Data

3 articles · Updated · PBS NewsHour · Jul 25

Summary

  • Three labels—closed, open-source and open-weight—describe how much of an AI model’s internals are public and reusable, not any “personality” of systems like ChatGPT.
  • Open-source software traditionally requires source code and the four freedoms to run, study, modify and redistribute it under licenses such as GPL, Apache, MIT and BSD.
  • Meta’s LLaMa, released on Feb. 24, 2023, published inference code and model weights, but the Open Source Initiative says limits on commercial reuse mean it is not truly open source.
  • Open-weight models such as DeepSeek and Alibaba’s Qwen have spread quickly because their reuse terms are looser, even though they still stop short of full open-source status.
  • The broadest definition of open-source AI also demands access to training data, a standard developers say is hard to meet because the datasets are enormous.

Insights

Are tech giants labeling AI models as open merely to dodge strict regulations while keeping the most valuable data locked away?
If open-weight AI models hide their training data, what invisible biases or stolen assets are we unwittingly building our future upon?
Could the push for fully open AI training data accidentally hand malicious actors the ultimate blueprint to exploit digital infrastructure?