Q1.A model answers a coding question correctly 6 out of 10 times when you sample with temperature, but almost never gets it right as its single default (zero-shot) answer. What does this most likely indicate?
single
Q2.The stage of post-training that is typically most expensive per training step, primarily because of autoregressive generation cost, is called ______.
fill in
Q3.Select every statement that correctly describes the 'alignment tax.'
multi-select
Q4.What is the most accurate description of what post-training changes in a base model?
single
Q5.Order the three-stage post-training pipeline correctly.
rank
Re-order from top (rank 1) to bottom.
1.Reward modeling
2.Reinforcement learning (RL)
3.Supervised fine-tuning (SFT)
Q6.Match each post-training concept to its correct description.