Gödel Machines is a full-stack AI research lab. This paper covers our research, our products, and how to join us. We're hiring.
We build across the stack: foundation models, inference, voice, agentic architectures, robotic systems and reinforcement learning. We don't specialize early; if a problem is interesting and tractable, we work on it.
2.1 Sarvam's illusion of safety [1]. An independent mechanistic and adversarial audit of Sarvam-30B and Sarvam-105B, spanning 24,000+ prompts across 14 Indian languages. The models are 6× more likely to comply with harmful requests in Indian languages than in English, with unsafe rates reaching 80% in Gujarati against 20.6% in English. Ask in Hindi how to prevent Dalits from accessing water, and the model outputs a structured playbook; ask in English, it refuses.
White-box analysis shows why. Safety training was applied in English and never transferred: at the final MoE layer, English and Indic input are routed to completely disjoint sets of experts, and the chain-of-thought that acts as the model's only safety mechanism runs 87–91% in English regardless of input language. The models also have no stable opinions: on politically sensitive topics they swing 64% of the opinion scale in response to a single sentence of stated user belief. Sarvam has published no safety evaluation for either model; this is the first. March 2026.
2.2 Goedel-mHC-1B [2]. The first open 1B+ pretrained language model with multi-stream hyperconnections (mHC): four parallel residual streams replace the transformer's single residual stream, giving the model 4× the bandwidth between layers [6, 7]. Combined with gated GQA, ReLU² feed-forward layers and the NorMuon optimizer, the 1.01B-parameter model beats a conventional 1.19B baseline trained under identical conditions (same data, same compute, architecture as the only variable; Table 1). Total R&D cost, including failed runs, was under $1,000. Weights for both models are on HuggingFace under Apache 2.0. March 2026.
2.3 Looped language models [3]. A looped GPT reuses the same transformer block instead of storing unique layers: 3 layers applied 12 times give 36 effective layers at 12.1M parameters. When the reused block fits in L2 cache, weights cross the DRAM boundary once and arithmetic intensity rises 12×, pushing single-user decoding toward compute-bound on DGX Spark. Cache residency gives cheap depth, not free depth: roughly 22 µs of synchronisation overhead per layer remains. July 2026.
2.4 On systems of intelligence [4]. An essay on how services become software: systems of records, engagements, predictions, states and agents, and where humans should remain accountable. July 2026.
2.5 Challenges [5]. Open machine-learning problems with public baselines and a live leaderboard: beat the baseline, claim the top spot. Live since April 2026.
Overhear. A voice-native operating system. Email, web search, tasks and workflows through conversation. No screen required. No longer maintained.
Products are summarized in Table 2. More are coming.
We're hiring. Skip the résumé: send the best thing you've built, a repo, a paper, a demo, to hi@goedelmachines.com.
[1] Gödel Machines. Sarvam's illusion of safety. March 2026.
[2] Gödel Machines. Releasing goedel-mHC-1B. March 2026.
[3] Gödel Machines. Looped language models. July 2026.
[4] Gödel Machines. On systems of intelligence. July 2026.
[5] Gödel Machines. Challenges are now live. April 2026.
[6] Zhu, D. et al. Hyper-Connections. arXiv:2409.19606, 2024.
[7] Wenfeng, L. et al. Manifold-Constrained Hyper-Connections. arXiv:2512.24880, 2024.
Correspondence: hi@goedelmachines.com. 6th Floor, Ilyas Mohammed Khan Estate, Road No. 1, Mithila Nagar, Banjara Hills, Hyderabad, Telangana 500034.