Discussion about this post

User's avatar
Leto II's avatar

Waiting your post about comparing Laguna S 2.1 with other models in similar size, Trust me bro benchmarks are nothing compared to your ones )

Thanks.

Will Hampson's avatar

I haven't gotten very good results with any Laguna model. I ran Laguna S 2.1 on DeepSWE and it didn't even achieve a fraction of the claimed bench. It was then released they changed parameters of the benchmark to 6 hours per task, among other things. Also claim it only performs like that in their pool harness, not the DeepSWE default. Their pool harness isn't open source so who knows what is happening in their to achieve the score, but i have been told by someone who looked at the binary there is a hidden advisor feature in there to call a smarter model that is not in the public documentation of the pool harness. maybe that's just for RL training or maybe that explains the score.

No posts

Ready for more?