the 8 hours vs 2 weeks framing is the part i'd want more on. generating
candidates got cheap, checking them didn't. what does the funnel actually look
like ,of the candidates from an 8 hour run, how many make it to synthesis?
asking because i hit the same shape in a much dumber domain and what got me was that the failures were quiet. nothing errored, output looked normal, it was just wrong in a way only someone who knew the domain would catch.
You can see from our the benchmark that only one of the candidates proposed was determined to be worth synthesizing. Each individual candidate generation is quick, 8 hours is required for the model to iterate with various tools to find ones worth submitting.
We found that speaking to domain experts was critical in desigining a rubric that could catch these silent synthesis recipe failures, before we attempt the longer 2 week synthesis effort.