Project Info

WrongWorlds

Devpost

Your wrong answer becomes a world you must investigate, fracture, rebuild, and escape. ๐ŸŽฎ Play Judge Mode โ€ข ๐ŸŒ Live Website โ€ข ๐Ÿ’ป GitHub โ€ข ๐Ÿ”ฌ Evidence Archive ๐Ÿ’ก

Inspiration

Most educational AI systems respond to a misconception by explaining why the learner is wrong. That helpsโ€”but it does not always change the mental model that produced the mistake. A learner may repeat the corrected answer while continuing to reason in exactly the same way. WrongWorlds began with a different question: What if a learner could enter a world where their wrong answer was actually true? Instead of immediately correcting the learner, WrongWorlds transforms the misconception into an explorable reality. Inside that reality, the learner must: ๐Ÿ” Investigate the world. ๐Ÿงฉ Collect contradictory evidence. ๐Ÿ’ฅ Fracture its false laws. ๐ŸŒ  Rebuild a stronger explanation. ๐Ÿšช Prove that the new reasoning transfers to another context. The result is the Museum of Possible Realities: an educational game concept where every misconception can become a world with its own rules, evidence, experiments, and escape condition. ๐ŸŽฎ What It Does WrongWorlds is a five-act illustrated reasoning game. A learner begins with a belief: โ€œThe marketing campaign caused the sales increase.โ€ GPT-5.6 interprets the learnerโ€™s language and converts it into a structured belief model. The learner then enters Signal Cityโ€”a world governed by that explanation. Rather than receiving the correct answer immediately, the learner must explore what their belief predicts, test competing explanations, and reconstruct the causal model using evidence. ๐Ÿงญ The Five-Act Experience ๐Ÿ“Š The Central Causal Reveal $$ +70 = +10_{\text{campaign}} + 40_{\text{festival}} + 20_{\text{district}} $$ The campaign contributed to the increaseโ€”but it was not the only cause. โœจ Why It Is Different WrongWorlds does not treat a misconception as a sentence to correct. It treats the misconception as a world to test. The learner must: ๐ŸŒ Enter the reality created by the belief. ๐Ÿ‘๏ธ Observe what that belief predicts. ๐Ÿงช Run controlled comparisons. ๐Ÿงฉ Collect evidence. ๐Ÿ’ฅ Watch the false reality fracture. ๐ŸŒ  Rebuild a stronger causal model. ๐Ÿš‰ Transfer the reasoning to another world. The learner does not escape by receiving the answer. The learner escapes by demonstrating better reasoning. ๐Ÿ—๏ธ How We Built It WrongWorlds was built with Codex and GPT-5.6, using a strict separation between: ๐Ÿค– Generative interpretation โš™๏ธ Deterministic educational authority This separation allows the system to understand flexible learner language without allowing the model to invent evidence, change simulation outcomes, or manipulate scores. ๐Ÿค– GPT-5.6 Interprets Learner Language GPT-5.6 converts free-form learner language into a structured belief model. It identifies elements such as: The proposed cause. The claimed effect. The learnerโ€™s reasoning patterns. Relevant variables. Confidence levels. GPT-5.6 is responsible for understanding what the learner believesโ€”not for deciding whether that belief succeeds inside the simulation. โš™๏ธ Deterministic Engines Govern the World GPT-5.6 does not determine: Evidence. Simulation results. State transitions. Fracture eligibility. Scores. Transfer outcomes. Those responsibilities remain with deterministic TypeScript domain engines controlling: ๐Ÿ™๏ธ Signal City experiments. ๐Ÿงฉ Evidence generation. ๐Ÿ’ฅ Fracture eligibility. ๐Ÿ“Š Causal contribution values. ๐ŸŒŒ The Reasoning Change Map. ๐Ÿš‰ Bloom Station transfer. ๐Ÿ Completion and scoring. ๐Ÿงฑ System Architecture ๐Ÿ” Core Design Principle GPT-5.6 interprets the learner. Deterministic systems retain authority over the world. ๐Ÿค Built With Codex Codex was the primary development partner throughout the project. It supported: ๐Ÿ›๏ธ Architecture and domain modeling. ๐Ÿ’ป Implementation. ๐Ÿงช Automated testing. ๐Ÿ” Deterministic replay. โ™ฟ Accessibility. ๐ŸŽจ Visual integration. ๐Ÿ–ผ๏ธ Asset validation. ๐Ÿ“ Documentation. ๐Ÿš€ GitHub and Vercel release workflows. ๐Ÿ” Evidence integrity checks. ๐Ÿ“ฆ The Final Application Includes A complete five-act illustrated Gamefront. English and Spanish localization. Keyboard navigation. Reduced-motion support. Accessible 2D parity. Responsive mobile layouts. Deterministic replay. Reproducible judge-facing screenshots. A public Evidence Archive. ๐Ÿงฑ Challenges We Ran Into 1. ๐Ÿง  Separating Interpretation From Authority The hardest architectural decision was determining exactly where GPT-5.6 should stop. We wanted the model to understand natural learner language, but we did not want it to: Invent evidence. Alter simulation outcomes. Influence scoring. Change state transitions. Solving this required explicit contracts between the belief interpreter and the deterministic domain engines. 2. ๐ŸŽฎ Turning an Editorial Prototype Into a Game The first Judge Mode was technically complete, but it felt more like an educational website than a world a learner could truly enter. We rebuilt the experience around: ๐Ÿ‘ค A recurring protagonist named Echo. ๐ŸŒ Two visually distinct worlds. ๐ŸŒ€ Portals. ๐Ÿงฉ Evidence objects. ๐Ÿ“ Causal hotspots. ๐Ÿ’ฅ A cinematic reality-fracture moment. ๐ŸŒŒ A puzzle-driven rebuilding sequence. ๐Ÿšช A final escape state. The visual transformation changed how the experience felt without modifying the underlying causal engine or frozen evaluation artifacts. 3. ๐Ÿ” Maintaining Deterministic Replay The same replay needed to produce the same: State. Evidence. Visuals. Screenshots. Final result. Every single time. The final screenshot pipeline generated byte-identical public assets across repeated runs. 4. ๐ŸŽจ Cleaning and Integrating Generated Assets The original game assets contained chroma-key edge contamination. We created a deterministic cleanup process that: Preserved image dimensions. Preserved alpha transparency. Removed green spill. Retained fine hair details. Preserved crystals and energy glows. Regenerated the public visual package from the real product. 5. โ™ฟ Preserving Accessibility Without Losing the Game Feeling WrongWorlds supports: โŒจ๏ธ Keyboard interaction. ๐Ÿง˜ Reduced motion. ๐Ÿ—บ๏ธ Accessible 2D mode. ๐ŸŒŽ English and Spanish. ๐ŸŽฏ Visible focus states. ๐Ÿ“ฑ Mobile layouts. The challenge was preserving the causal revelation for learners who cannotโ€”or prefer not toโ€”use cinematic motion. ๐Ÿ† Accomplishments We Are Proud Of โœ… Built a complete five-act educational game experience. โœ… Created two visually distinct reasoning worlds. โœ… Preserved deterministic authority over evidence and scoring. โœ… Implemented three real causal experiments in Signal City. โœ… Added transfer assessment through Bloom Station. โœ… Built deterministic replay with zero application POST requests. โœ… Deployed a Git-backed production release on Vercel. โœ… Published a public GitHub repository. โœ… Published a public Evidence Archive. โœ… Added English and Spanish localization. โœ… Added mobile, keyboard, and reduced-motion support. โœ… Added Accessible 2D parity. โœ… Generated reproducible README, Devpost, Open Graph, and video assets. ๐Ÿงช Final Verification ๐Ÿ“š What We Learned ๐Ÿง  A Correction Is Not the Same as a Changed Mental Model A learner can remember the right answer without understanding why the original reasoning failed. WrongWorlds makes the misconception: Visible. Testable. Explorable. Breakable. Rebuildable. ๐Ÿ”„ Transfer Matters A learner has not fully learned a reasoning pattern if they can only repeat it in the original example. Bloom Station exists to test whether controlled-comparison reasoning transfers to a completely different context. ๐Ÿ” AI Becomes Stronger When Its Authority Is Bounded GPT-5.6 is valuable because it can interpret the many ways a learner may express a belief. The deterministic engine is valuable because it guarantees that: Evidence remains stable. Scores remain reproducible. Simulations remain auditable. Evaluation remains trustworthy. The strength of the system comes from giving each component a clearly defined responsibility. ๐ŸŽจ Visual Design Changes How Technical Work Is Understood The original implementation already contained the causal engine, but the illustrated Gamefront made the educational concept immediately understandable. The fracture scene became the clearest expression of the product: One apparent cause breaking into multiple contributions. ๐Ÿš€ What Is Next WrongWorlds currently contains: ๐Ÿ™๏ธ One primary exhibit: Signal City. ๐Ÿš‰ One transfer world: Bloom Station. The Museum of Possible Realities could expand into new exhibits covering: ๐Ÿ”ฌ Science misconceptions. โž— Mathematical reasoning. ๐Ÿ’ฐ Economics. ๐Ÿ“œ History. ๐Ÿ“Š Statistics. โš™๏ธ Systems thinking. ๐Ÿ“ฐ Media literacy. ๐Ÿ”ฎ Future Work Could Include Teacher-authored WrongWorlds. Classroom cohorts. Learner progress histories. Adaptive transfer worlds. Controlled educational studies. Long-term learning and retention measurement. โš ๏ธ Current Limitation The current Evidence Archive is exploratory and is not presented as proof of educational efficacy. Larger learner studies would be required to measure: Long-term learning. Retention. Reasoning improvement. Transfer across contexts. ๐ŸŽฎ Try WrongWorlds ๐Ÿ› ๏ธ Built With Codex ยท GPT-5.6 ยท TypeScript ยท React ยท Next.js ยท Node.js ยท Playwright ยท Vitest ยท Vercel ยท GitHub ยท CSS Modules ยท HTML5 ยท WebP ยท Accessibility ยท Internationalization ยท Deterministic Simulation ๐ŸŒŒ Enter the wrong world. Test its laws. Rebuild the truth. Escape with better reasoning.

Analysis

Compare with all teams

View

Metric

Figures cover GitHub contributors during the hackathon window. A co-authored commit counts in full for each author, so per-member totals add up to more than the whole-team figures.

Technology

Found in codeClaimed only
  • CSSIn code
  • Next.jsIn code
  • OpenAIIn code
  • ReactIn code
  • TypeScriptIn code
  • VercelClaimed

5 of 6 appear in the indexed code. 1 claimed on Devpost could not be matched to code, which may simply mean the tool leaves no trace in the repository.

AI coding agents

  • CodexConfig

Detected from committed agent config files and commit authorship. Absence of a signal is not proof an agent was unused.

Codebase size

Source size

1004 KB

Source files

182

Counts recognized source files only; vendored directories, binaries and lockfiles are excluded, so this is smaller than the repository on disk.

0 stars