AI can generate a stack of research proposals before a human finishes coffee. That is not the same thing as knowing which proposal is worth six months of work.
Anthropic researchers introduced TASTE, a benchmark that measures how closely models agree with experienced AI safety researchers when choosing between pairs of research proposals. The dataset contains 92 comparisons and estimates human agreement with its labels at 77 percent. The strongest model result reported was 60 percent.
The benchmark targets judgment rather than output fluency. Many research questions do not have a clean answer key. A proposal can be clever but irrelevant, important but impossible, or technically elegant while aimed at the wrong failure mode. Those distinctions are where taste, experience, and domain knowledge become operational assets.
The team built the benchmark from model-generated proposals inspired by 93 human-written ideas. Researchers scored proposals individually, discussed disagreements in pairs, then revised their preferences. Filtering for strong-confidence judgments and larger score gaps produced the final comparison set. That process matters because human judgment is noisy too.
The result should not be inflated into a universal verdict on machine creativity. Ninety-two pairs is a small benchmark, confidence intervals are wide, and the labels approximate researcher agreement rather than objective truth. Anthropic explicitly says larger and more diverse evaluations are needed.
Still, the gap exposes a useful boundary. Models can accelerate the production of options faster than organizations can improve the quality of selection. More ideas do not create more progress when the evaluation layer remains weak.
Builders chasing automated research need a system that separates proposal generation from proposal authority. Use models to expand the search space, pressure-test assumptions, and surface neglected approaches. Keep expert review, disagreement, and confidence visible until the machine proves it can recognize a strong direction without being hypnotized by polished language.
LaunchPad positionAutomating research is not only a generation problem. The harder leverage point is deciding which question deserves scarce time, compute, and human attention.
This report draws on the linked primary sources and reputable reporting. Company statements are treated as claims until independently demonstrated.
