Equation 9 · Comparing the Main Approaches to Training Data and Synthetic Data
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
the fraction. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
so as a task gets harder and p falls, rejection sampling gets steadily more expensive for the same yield — precisely the pattern that shows up in Llama 2’s own pipeline as a second mechanism being layered on once sampling alone stopped paying for itself.
Sources cited in the article section
- [6] Llama 2: Open Foundation and Fine-Tuned Chat Models ↗
- [7] Scaling Relationship on Learning Mathematical Reasoning with Large Language Models ↗
- [9] Constitutional AI: Harmlessness from AI Feedback ↗
- [3] Self-Rewarding Language Models ↗
- [13] Textbooks Are All You Need ↗
These citations give research context. Read each source to check which claims it supports.
Return to Comparing the Main Approaches to Training Data and Synthetic Data