The Current Gap

On standard benchmarks like MMLU and HumanEval, open-source models are catching up fast. But in complex reasoning, instruction following, and safety, closed-source models still hold advantages.

Catch-Up Paths

  • Data quality: Open-source training data quality is improving, especially code and math
  • Synthetic data: GPT-4-generated data widely used for training open models
  • Community innovation: LoRA, quantization, fine-tuning techniques leverage closed-source outputs

At current pace, open-source models will match closed-source in most conventional tasks within 12-18 months. Frontier reasoning may take longer. The models mentioned in China’s LLM Competitive Landscape are key forces in this catch-up.