Show HN: WorldBuild Bench repo: testing LLM world coherence with 3D games
I built WorldBuild Bench because, as we all know, llm bench scores often say something very different from what models actually feel like to use. It's really dependent on the type of tasks. I personally want to test spatial, temporal, and causal coherence in an interactive 3D world. Does the model understand where things are, world stays consistent over time and do the consequences make sense? There is a million people generating random games here and there on yt, but I want something that I can















