I quite disagree
Paired T tests in a real environment provides a solid empirical evidence.
You could quite easily run a model with identical hands as well
Perhaps on snowie or other similar tool
Later adds...
Don't get me wrong, I don't disagree with the OP's premise.
I have been bluffing all week, so I can provide observational and anecdotal evidence
In the real world lots of people are building better bots and the best way to improve them is testing in the real world.
It seems to me that you are talking about running a bot who is playing vs other bots and letting 2 bots play identical set of hands? how is that going to help you to realize if you 3 barrel bluffs are profitable vs human population in real poker?
In the real world lots of people are building better bots and the best way to improve them is testing in the real world.
So you are building bots ?