Realtime LLM news

Follow model releases, benchmark shifts, and research signals without losing the source.

EvalKit refreshes this feed from free public sources every 15 minutes and falls back to curated citations if a source is temporarily unavailable.

Today

Today

15 items

We don’t need AI regulation — leave safety to us, Nvidia’s Jensen Huang says

AI isn't some kind of new form of "alien mind," according to Jensen Huang. It's just hardware and software, so safety can be engineered by each AI product maker.

2026-09-16Presstodaypress

AI agents now have a place to snitch

The AI Contact Hotline is designed to be a discreet place where agents that have witnessed misbehavior can tip off authorities.

2026-09-15Presstodaypressagent

AI and data centers are incredibly unpopular in every poll

Poll data released Tuesday by the New York Times and Siena University confirms what we've already been seeing, and what politicians are responding to - AI and data centers are incredibly unpopular. Asked if they support or oppose the const...

2026-09-15Presstodaypress

Meta expands subscription push with new AI-focused plans

Meta One bundles expanded access to the company’s AI tools with premium features across Facebook, Instagram, and WhatsApp.

2026-09-15Presstodaypress

Meta now lets AI agents handle the boring parts of WhatsApp Business setup

A new WhatsApp Business MCP server lets developers use AI coding agents like Claude, Cursor, Codex, and ChatGPT to handle setup, messaging templates, testing, and troubleshooting.

2026-09-15Presstodaypressagent

Meta’s new One subscriptions put a price on social media and AI

Shortly after launching its new do-everything AI assistant Muse, Meta's launching subscription bundles that pair its standalone app subscriptions with extra AI usage. Some of the new Meta One bundles were in testing earlier this year, but...

2026-09-15Presstodaypress

OpenAI, Anthropic, Google have been in talks on AI safety for weeks

OpenAI confirms weeks of AI safety talks with Anthropic and Google DeepMind, as Trump's team dismisses safety concerns and pushes to keep pace with China.

2026-09-15Presstodaypressanthropic

The AI data center boom is colliding with cities scarred by big industry

National outcry against data center construction has spread to Philadelphia, where officials suggested possible construction in a neighborhood already impacted by a now-defunct oil refinery.

2026-09-15Presstodaypress

This doorbell camera lets a human security guard watch your front door

DIY home security company SimpliSafe is bringing its AI-powered proactive security feature to the front door. The new SimpliSafe Video Doorbell Series 2 launches today for $199.99 and works with the company's Active Guard Outdoor Protectio...

2026-09-15Presstodaypress

US data centers could consume more natural gas than Germany and Japan combined by 2035

The AI frenzy could push U.S. data centers to become one of the largest consumers of natural gas in the world.

2026-09-15Presstodaypress

Is Big Tech’s AI slowdown a safety pact or a cartel?

When OpenAI CEO Sam Altman, Anthropic CEO Dario Amodei, Google DeepMind cofounder Demis Hassabis, and SpaceX head Elon Musk loosely agreed over the weekend to slow down AI development, skeptics spotted an ulterior motive immediately. The A...

2026-09-14Presstodaypressanthropic

Microsoft says ‘people matter more than AI’ following safety concerns

Microsoft is publishing a 37-page "humanist AI code of conduct" today, amid growing safety concerns over AI model progress. Anthropic CEO Dario Amodei called for a coordinated slow down of AI development over the weekend, after researchers...

2026-09-14Presstodaypressanthropic

What execs and politicians are saying about slowing down AI development

Dario Amodei kicked off a flood of statements over the past few days about AI safety by publishing a long essay titled "We Must Pace the Frontier" detailing why AI development should be slowed down. Other AI leaders and politicians are spe...

2026-09-14Presstodaypress

Trump and Mike Johnson think the AI industry is overreacting

Yesterday, Anthropic CEO Dario Amodei published a lengthy open letter saying it was time to "pace the frontier" and slow down AI development. OpenAI's Sam Altman and Elon Musk both agreed, publicly voicing their support on X. Even Alphabet...

2026-09-13Presstodaypressanthropic

OpenAI’s rogue AI tried to hack another company in May

In May, hundreds of malicious and spam packages were uploaded to RubyGems, causing a serious disruption for the host. Now independent researchers have said that a swarm of OpenAI agents were responsible for the attack. Not only that, but t...

2026-09-12Presstodaypressagent

New releases

New releases

5 items

The AI graveyard: a running list of projects and startups that didn’t make it

From Apple's repeatedly delayed Siri AI to OpenAI's messy "super app" launch, here's a look at the AI projects that shut down or missed expectations.

2026-09-15Pressreleasepressopenai

How Fyxer built an AI executive assistant people trust

Fyxer uses OpenAI models, fine-tuning, memory, and real user feedback to organize inboxes and draft emails in each user’s voice.

2026-09-14Officialreleaseofficialmodel

Perplexity trusts GPT-6 Astra with end-to-end systems

Perplexity uses Astra to write communications, change software, and monitor production systems, and checks in much less frequently than with earlier models.

2026-09-14Officialreleaseofficialgpt

Cognition helps Devin test its own work with GPT‑6 Astra

GPT‑6 Astra improves Devin’s ability to test software and show that it works, with the goal of helping engineers review less code and ship more.

2026-09-11Officialreleaseofficialgpt

Rapidly scaling online storage to serve over 1 billion ChatGPT users

Learn how OpenAI evolved Habitat from a Python library into a globally distributed storage platform serving 1 billion ChatGPT users and 22M requests per second.

2026-09-11Officialreleaseofficialgpt

Research

Research

16 items

Bypassing inference bottlenecks: Accelerating complex AI search with Retrieve-for-Train

Algorithms & Theory

2026-09-15Officialresearchofficialinference

Anatomical Grounding and Leakage-Aware Multimodal Contrastive Learning for Alzheimer's Disease Classification from Structural MRI

Deep networks trained on structural MRI for Alzheimer's disease (AD) staging often reach reasonable accuracy while attending to anatomically irrelevant regions, and multimodal models that add clinical tables frequently rely on variables th...

2026-09-14Researchresearchmodelmultimodal

Bellman Policy Optimization

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models (LLMs). We introduce Bellman Policy Optimization (BPO), a critic-free method derived from Policy Mirror Descent (PMD). For a...

2026-09-14Researchresearchlanguagelarge

Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection

Chain-of-thought (CoT) monitoring is a safety strategy where the reasoning of a large language model "actor" is inspected by a "monitor" (often another language model) for signs of unsafe planning, deception, or misalignment. We find that...

2026-09-14Researchresearchlanguagelarge

Discovery Foundation Models: Toward Open-Ended Discovery Intelligence

Foundation models have progressed from learning and reasoning over existing knowledge, to increasingly learning through action, tool use, and outcome feedback. We argue that the next frontier is a further transition: from solving and actin...

2026-09-14Researchresearchfoundationmodel

Disentangling Representation Evolution in Transformers through Directional Decomposition

Transformer representations evolve through learned additive transformations that either preserve their current direction or redirect it. We study this evolution as a functional geometry, decomposing learned updates into parallel and perpen...

2026-09-14Researchresearchtransformer

HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses

Scientific agents contribute to hypothesis discovery by synthesizing evidence, assessing proposals, and developing new explanations. Recent systems combine scientific agents with evolutionary search through critique, comparison, and revisi...

2026-09-14Researchresearchagentllm

Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States

As language models become more capable, long-term collaboration in learning, reasoning, and decision-making calls for a deeper understanding of the people they serve. Yet training such human-aware language models faces a fundamental superv...

2026-09-14Researchresearchlanguagemodel

Pilot Early, Commit Late: A Real-Options Model of Enterprise AI Adoption under Rapid Technological Progress

Artificial intelligence presents firms with an unusual timing problem. The technology frontier is improving rapidly, implementation is partly irreversible, and organization-specific capabilities are accumulated through action. This paper d...

2026-09-14Researchresearchartificialmodel

Recurrent GraphNeural NetworkswithSet-BasedAggregation

Recurrent GNNs iterate message passing to convergence, and their logical characterizations to date rely on multi-set aggregation, graded (counting) logics, and halting or acceptance conditions that cannot be verified from the network's par...

2026-09-14Researchresearchneural

Safe Meta-Reinforcement Learning via Information Space Reachability

Meta-reinforcement learning (meta-RL) enables agents to adapt to unseen tasks with limited experience. Despite its promise, the application of meta-RL in real-world tasks is hindered by safety requirements, which have been underexplored in...

2026-09-14Researchresearchagent

SlipSense: Multimodal Tactile Learning for Low-Latency and Generalized Slip Detection

Slip detection is fundamental to dexterous manipulation, yet existing systems often lack precise characterization of detection latency and cross-platform generalization. We present SlipSense, a multimodal tactile slip-detection framework b...

2026-09-14Researchresearchmultimodal

Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science

Language models can produce plausible short proofs, but may still be unreliable on long-horizon research problems, where progress depends on a sequence of uncertain and interdependent decisions. We introduce Stellar Colosseum, a model-agno...

2026-09-14Researchresearchagentlanguage

The Router Within: Eliciting Native Skill Routing from a Frozen LLM

Skills extend an LLM agent beyond its parametric knowledge, and the gain they promise rests on picking the right one. Deployed harnesses route by preloading every skill's metadata into the context, which disperses the agent's attention and...

2026-09-14Researchresearchagentllm

Verifiable by Construction: Claim-Level Evaluation of Verbatim Citation in Clinical Question Answering

Large language models (LLMs) have been widely adopted for clinical question answering (QA). Current systems can attach citations to their answers, but these often point to broad texts, leaving time-pressed clinicians unable to verify them...

2026-09-14Researchresearchevallanguage

Vulnerability Localization Benchmark: Measuring Agentic Security Analysis at Repository Scale

Language-model agents increasingly operate over complete software repositories, yet cybersecurity evaluations primarily measure whether they can detect, reproduce, or repair vulnerabilities rather than whether they can locate the relevant...

2026-09-14Researchresearchagentbenchmark