← 返回案例流
CASE / reddit-1wlm50lSOURCE / Reddit
Reddit / REDDIT
SOURCE / REDDIT
SSilver-Champion-4846@Silver-Champion-4846

The Scaling Lesson. What are we missing?

Consider the difference between GPT 1 and GPT 6. They are both decoder-only language models, yet there's so much difference between them. Data scale, data quality, sft, rl, etc. But the important lesson, I think, is that the preliminary capabilities of a specific manifestation of a machine learning architecture should not be viewed as the absolute limit of that architecture's potential, seeing as we went from 'autocomplete on steroids...
中文翻译

中文参考整理中,欢迎补充。

In progress; contributions welcome.
查看 Reddit 原链接

来源说明

本条为社区整理,仅展示公开来源与链接。代码、文本、头像及商标权利归原作者和原平台所有;Laya 模型归 Convai Innovations 所有。

整理方式:自动采集 · 整理时间:2026年9月22日 14:30