/AI Weekly/Issue 145

Issue #145 16 stories

Global AI Weekly

Mythos could fix AI hallucinations

Published Tuesday, April 14, 2026

In this issue

Highlights

3 stories
Mythos could fix AI hallucinations
red.anthropic.com

Mythos could fix AI hallucinations

Anthropic’s Mythos preview goes after the real problem behind hallucinations: AI does not just get things wrong, it builds convincing stories. This work explores how models handle truth, uncertainty, and explanation, and why accuracy alone is not enough. The idea is simple but powerful: make AI show its doubt instead of hiding it. If this direction sticks, it could reshape how we trust everything AI says.

The Role of Tech and AI in the Artemis II Moon Mission
aimagazine.com

The Role of Tech and AI in the Artemis II Moon Mission

NASA’s Artemis II mission is gearing up to launch the Space Launch System (SLS) rocket alongside the Orion spacecraft, setting the stage for groundbreaking advancements in lunar exploration. Equipped with cutting-edge technology and artificial intelligence, the mission aims to enhance crew safety, improve navigation, and pave the way for future deep space travel. This effort represents a pivotal step in returning humans to the Moon and expanding our reach into the cosmos.

In this issue

Research

3 stories
Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents
huggingface.co

Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents

This paper introduces Claw-Eval, a novel approach aimed at creating more reliable and trustworthy evaluations for autonomous agents. By addressing current challenges in assessing these systems, it proposes a framework that ensures consistency and fairness while highlighting key areas for improvement. The work emphasizes the importance of robust benchmarks to support the growth and deployment of AI-driven agents.

Embarrassingly Simple Self-Distillation Improves Code Generation
huggingface.co

Embarrassingly Simple Self-Distillation Improves Code Generation

This paper explores how an embarrassingly simple self-distillation approach can enhance code generation performance. By iteratively refining a model with its own predictions, this method achieves notable improvements without the need for complex processes or additional supervision. The study highlights the effectiveness of simplicity in advancing code generation tasks.

In this issue

Video

2 stories
Opinionated agentic development and sharing
youtube.com

Opinionated agentic development and sharing

In Episode 3 of the Made for Dev @DockerInc special, Sammy Deprez and Oleg Šelajev explore the capabilities of Docker Agent, a tool for building custom, portable AI agents. Oleg showcases the "Agent-as-Code" philosophy, demonstrating how to configure an agent's personality, AI model, and toolset using a simple YAML file. The episode highlights features like running agents locally with the Docker Agent CLI and integrating tools to handle tasks such as research, coding, and debugging. If you've been curious about creating shareable AI assistants tailored to your needs, this episode has you covered.

Claude Mythos, Project Glasswing and AI cybersecurity risks
youtube.com

Claude Mythos, Project Glasswing and AI cybersecurity risks

This week's Mixture of Experts podcast explores key AI topics, including Anthropic's decision to withhold its Mythos model and the implications for AI security, along with a breakdown of the financial strategies of OpenAI and Anthropic as they tackle different market sectors. The discussion also covers whether AI can rediscover historic scientific breakthroughs, featuring findings from the GPT-1900 model experiment. To wrap up, IBM Fellow Aaron Baughman showcases the Masters Vault, an AI innovation that enables seamless searching of decades worth of Masters golf footage using natural language.

In this issue

Articles

3 stories
VS Code Just Turned AI Agents Into Your New Dev Team
code.visualstudio.com

VS Code Just Turned AI Agents Into Your New Dev Team

The latest Visual Studio Code update quietly levels up AI from autocomplete to full agents that can plan, act, and iterate across your workspace. These agents go beyond suggestions, handling multi-step tasks, tool use, and context-aware changes inside your codebase. It marks a shift from passive assistance to active collaboration, where AI can actually execute work. If you thought Copilot was useful, this release hints at what happens when it starts behaving like a real teammate.

glm-5.1:cloud
ollama.com

glm-5.1:cloud

GLM-5.1 is a cutting-edge model designed for advanced agentic engineering, offering vastly improved coding capabilities over its predecessor. It sets a new benchmark by excelling in SWE-Bench Pro and outperforms the earlier GLM-5 model with a significant lead, showcasing its exceptional performance and innovation.

In this issue

Upcoming Events

1 story
AgentCamp - Coming to a City Near You
globalai.community

AgentCamp - Coming to a City Near You

AgentCamp continues to grow as a global series of hands-on gatherings dedicated to building and experimenting with AI agents. These community-driven events bring developers, founders, and AI enthusiasts together for practical sessions, collaborative building, and open exchange of ideas. Hosted in cities around the world, AgentCamp focuses on real-world experimentation, giving participants the space to prototype agent workflows, explore emerging tools, and learn directly from peers working at the edge of autonomous AI. Join the community to build, share, and help advance what AI agents can do in practice.

In this issue

Code

3 stories
In this issue

Podcast

1 story
Machine Learning Street Talk
open.spotify.com

Machine Learning Street Talk

Machine Learning Street Talk (MLST) offers engaging conversations with leading experts in AI, covering topics like cognitive science, neuroscience, and the philosophy of mind. The show provides in-depth analysis of current developments in AI, emphasizing intellectual diversity while cutting through the hype. Hosted by Tim Scarfe, Ph.D., with regular contributions from MIT's Dr. Keith Duggar, MLST delivers a rigorous and wide-ranging perspective on the field.

Every Tuesday

Get the next issue in your inbox.

The AI links worth your time, curated by the community and always free.