You can see from our the benchmark that only one of the candidates proposed was determined to be worth synthesizing. Each individual candidate generation is quick, 8 hours is required for the model to iterate with various tools to find ones worth submitting.
We found that speaking to domain experts was critical in desigining a rubric that could catch these silent synthesis recipe failures, before we attempt the longer 2 week synthesis effort.