The Current Gap
On standard benchmarks like MMLU and HumanEval, open-source models are catching up fast. But in complex reasoning, instruction following, and safety, closed-source models still hold advantages.
Catch-Up Paths
- Data quality: Open-source training data quality is improving, especially code and math
- Synthetic data: GPT-4-generated data widely used for training open models
- Community innovation: LoRA, quantization, fine-tuning techniques leverage closed-source outputs
At current pace, open-source models will match closed-source in most conventional tasks within 12-18 months. Frontier reasoning may take longer. The models mentioned in China’s LLM Competitive Landscape are key forces in this catch-up.