/AI Weekly/Issue 168

Issue #168 12 stories

Global AI Weekly

OpenAI, Anthropic and Google Join Forces on Safety

Published Tuesday, September 22, 2026

In this issue

Highlights

2 stories
AI Rivals Unite: OpenAI, Anthropic and Google Join Forces on Safety
techcrunch.com

AI Rivals Unite: OpenAI, Anthropic and Google Join Forces on Safety

OpenAI, Anthropic, and Google DeepMind have reportedly been holding private talks for weeks to tackle the growing risks of advanced AI. The unlikely alliance comes amid calls for independent safety evaluations, industry-wide standards, and stronger oversight of frontier models. OpenAI is also backing legislation that would allow independent organizations to verify AI safety practices. But with antitrust concerns and competitive pressures looming, can these rivals truly cooperate when the stakes are higher than ever?

GPT-6 Astra Dominates Math, But Claude Still Has the Coding Edge!
epoch.ai

GPT-6 Astra Dominates Math, But Claude Still Has the Coding Edge!

OpenAI's GPT-6 Astra tops Epoch AI's Capabilities Index with a score of 166, but the real story lies beneath the headline numbers. Astra sets a new record in mathematical reasoning, while Anthropic's Claude Fable 5.1 scores higher on software engineering benchmarks. Interestingly, Astra's overall score dropped as additional coding evaluations became available. With overlapping confidence intervals and rapidly evolving benchmarks, the findings highlight why choosing the right AI model depends on more than a single leaderboard position.

In this issue

Research

2 stories
AI Built a Frontier Chip in Just 2 Weeks. That Changes Everything
alphaxiv.org

AI Built a Frontier Chip in Just 2 Weeks. That Changes Everything

What if AI could design the hardware it runs on? Redwood claims exactly that: a frontier AI accelerator conceived, verified, and deployed from scratch in under two weeks, with minimal human input beyond the high-level spec. Built for low-power, ultra-low-latency inference, it reportedly beats Jetson Orin Nano on projected efficiency by a wide margin. More strikingly, the system can rapidly rework and redeploy new designs, hinting at a future of self-improving AI hardware.

AI Coding Agents Are Getting Smarter, But Are Their Leaderboards Lying?
arxiv.org

AI Coding Agents Are Getting Smarter, But Are Their Leaderboards Lying?

Can you really trust coding-agent leaderboards? A new study examines 254 SWE-bench submissions and finds that small differences between leading agents often lack statistical significance. The researchers also reveal that changing an agent's scaffolding can affect performance more than switching between top-ranked models. With leading agents solving many of the same problems, the paper challenges how we compare AI coding tools and proposes a more rigorous approach to evaluating their capabilities.

In this issue

Video

2 stories
ZoomIt for Mac: How AI Ported a 20+ Year-Old Windows Tool
youtube.com

ZoomIt for Mac: How AI Ported a 20+ Year-Old Windows Tool

Microsoft's Mark Russinovich has brought the legendary Sysinternals ZoomIt to macOS, using GitHub Copilot to accomplish in just two days what once seemed too costly to attempt. In this video, he demonstrates how AI helped rewrite the Windows utility in Swift, complete with screen magnification, annotation, and recording features. Discover how agentic coding is making previously impractical software projects possible, and what this means for the future of cross-platform development.

Agents Don’t Know What They Don’t Know
youtube.com

Agents Don’t Know What They Don’t Know

AI coding agents can generate code at lightning speed, but they don't always know when they've made a mistake. In this talk, Rob Zuber explores why reliable agentic development depends on tight feedback loops rather than smarter models alone. From automated testing and stop hooks to CI pipelines and human review, discover how to give coding agents the tools to catch their own mistakes, improve code quality, and deliver software you can actually trust.

In this issue

Articles

2 stories
Hollywood's First AI Actress Goes on a Press Tour. What Could Possibly Go Wrong?
techcrunch.com

Hollywood's First AI Actress Goes on a Press Tour. What Could Possibly Go Wrong?

Meet Tilly Norwood, the AI-generated actress whose first press tour is turning into an unexpected comedy. During an interview with Piers Morgan and actor Tom Conti, Tilly struggled to answer basic questions about her upcoming movie before suddenly switching to Chinese mid-sentence. Her creators have made her available for 75 simultaneous interviews, but the resulting mishaps raise an amusing question: If AI can replace Hollywood actors, shouldn't it at least be able to survive a press junket?

What happens when you give a Fruit Fly Brain 12,716 Cybersecurity Alerts.
intezer.com

What happens when you give a Fruit Fly Brain 12,716 Cybersecurity Alerts.

First, internet experimenters made simulated fruit fly brains play video games. Now, cybersecurity company Intezer has gone a step further, using a digital model of a fly's nervous system to analyze 12,716 real security alerts. Surprisingly, the fly-inspired system performed reasonably well. Unfortunately for our six-legged cybersecurity expert, randomly rewiring its brain produced better results, and a simple machine learning model outperformed them both. A fascinating experiment that proves sometimes the simplest solution really does win.

In this issue

Upcoming Events

1 story
October 20–22: NVIDIA GTC Berlin
nvidia.com

October 20–22: NVIDIA GTC Berlin

Europe's biggest AI moment is happening this autumn. GTC Berlin brings together developers, researchers, and industry leaders to go deep on the full five-layer AI stack, from energy, chips, and infrastructure to open models and physical AI.

In this issue

Code

2 stories
Stop Guessing! 5 Prompt Tricks That Actually Improve AI Output
kdnuggets.com

Stop Guessing! 5 Prompt Tricks That Actually Improve AI Output

Better AI results don't always require a better model. Sometimes, your prompts just need an upgrade. This practical guide explores five proven optimization strategies, from structured outputs and role-based instructions to few-shot examples, targeted reasoning, and automated prompt testing. Using a messy meeting transcript as a real-world example, it demonstrates how small, measurable changes can improve accuracy, reduce hallucinations, and produce more reliable results. The key takeaway? Stop tweaking prompts blindly and start testing what actually works.

Estimators in Scikit-LLM: A KDnuggets Cheat Sheet
kdnuggets.com

Estimators in Scikit-LLM: A KDnuggets Cheat Sheet

What if you could plug LLMs directly into your existing machine learning pipelines? Scikit-LLM makes it possible by wrapping language models in the familiar scikit-learn API. This handy cheat sheet explores zero-shot and few-shot classification, text embeddings, translation, and seamless integration with pipelines and cross-validation. Discover how to combine traditional ML workflows with the power of LLMs, without reinventing your entire toolkit. Just keep an eye on those API costs!

In this issue

Podcast

1 story
Is Open-Source AI About to Change Everything?
open.spotify.com

Is Open-Source AI About to Change Everything?

Open-source AI is no longer just a playground for hobbyists. As local models become more capable and API prices fall, businesses are reconsidering their dependence on expensive proprietary models. In this episode, Jordan Wilson explores the shrinking gap between open and closed AI, the rise of always-on agents, and the growing importance of choosing the right model for each task. Could local AI become the smarter choice for your next project?

Every Tuesday

Get the next issue in your inbox.

The AI links worth your time, curated by the community and always free.