Benchmarking Learning from Label Proportions (LLP) is trickier than it looks, because different LLP variants demand different evaluation setups. This preprint proposes methods to generate variant-specific datasets and guidelines for benchmarking LLP algorithms fairly.