The most interesting thing here is the kernel optimization graph.
It look like all models were still improving, when they cut off the experiment.
It reminds me of a genetic algorithm. The graph is the same: long plateaus and then massive leaps.
The only difference between the models seems to be how quickly they arrive.