LIVE-Last scan updating-53 sources active-897 signals today-BENCHMARKSSWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?
Source-linked AI news and decision tools

Track the AI changes that matter to builders.

AI on Radar monitors official releases, model and API changes, GitHub momentum, papers, and developer tools, then publishes concise reporting with visible sources and practical calculators.

01
LLM API selectorFilter source-backed model routes by workload, context, capabilities, and budget.
02
API cost calculatorEstimate monthly input, output, cache-read, and per-request charges.
03
GPU / VRAM calculatorEstimate model-weight and optional architecture-aware KV-cache memory.

Latest on the radar

5 signalsView all articles
FeaturedBenchmarksRDR82

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

A new benchmark, SWE Refactor Bench, has been introduced to evaluate the ability of coding agents to perform complex, whole-repository software migrations. Existing benchmarks are insufficient as they do not verify if the migration actually occurred, allowing agents to pass tests by copying original code. This benchmark addresses that gap by assessing both migration completeness and behavioral correctness.

2 min - 20h agocoding agents
InfrastructureRDR80

OpenAI Unveils Jalapeño Inference Chip

OpenAI has introduced Jalapeño, a custom-designed inference chip. This new hardware aims to significantly improve the speed and power efficiency of AI inference tasks.

1 min - 20h agoAI inference
AI ToolsRDR79

Wire It, Run It, Deploy It: AI Workflows in Gradio

This guide from Hugging Face explores building and deploying AI workflows using Gradio. It covers the process from initial setup and execution to final deployment, offering practical insights for developers.

1 min - 20h agoGradio
Developer ToolsRDR82

Import Setup and Recent Work from Other Agents

OpenAI has updated its developer tools to allow importing setup and recent work from other agents. This feature aims to streamline the workflow for developers using OpenAI's agent-based systems.

2 min - 1d agoagents
BenchmarksRDR68

New Hugging Face Datasets for PDE LLM Evaluation

Three new Hugging Face datasets have been released for evaluating Large Language Models (LLMs) on Partial Differential Equation (PDE) problems. These datasets focus on free-generation tasks and include tabular and text modalities.

2 min - 1d agoLLM

Evergreen decision library

6 compare / 6 alternatives / 6 use cases / 4 changesOpen compare hub

Signal surfaces

Live across tracked layers

Topic intelligence

Graph

Cross-source topics rebuilt from clusters, entities, source mix, and article decisions.

  • New Gemma-2-9b-it Datasets for Math, Physics, and Chemistry on Hugging FaceResearch Papers - 8 linked signalsRDR61
  • penfever/nemotron-gym-competitive-coding-minimax-m27-131k-traces_chunk1 Dataset signal on Hugging FaceModels - 24 linked signalsRDR59
  • ​ Fine-tuning preview CLI commandSignals - 3 linked signalsRDR58
  • ​ Model revision validation statusSignals - 3 linked signalsRDR58
12 active topicsOpen

GitHub momentum

Top 25

Open-source attention measured by star velocity, freshness, and topic relevance, not raw stars.

  • santifer/career-opsJavaScript - 67,113 stars - +2779 7dRDR92
  • holaboss-ai/holaOSTypeScript - 9,611 stars - +2202 7dRDR87
  • opensandbox-group/OpenSandboxPython - 14,063 stars - +1570 7dRDR89
  • QwenLM/qwen-codeTypeScript - 27,219 stars - +227 7dRDR92
8,847 repos trackedOpen

Model watch

Sourced

Releases and capability changes from Hugging Face, OpenRouter, Artificial Analysis, and source APIs.

  • AtesiT/ru-llm-judge-datasetAtesiT - context not listed - unknownRDR60
  • aziz9788/saudi-llmaziz9788 - context not listed - unknownRDR58
  • orcarouter/spoken-multihop-ragorcarouter - context not listed - unknownRDR57
  • swarmagents/ufr-multiturn-agents-15k-previewswarmagents - context not listed - unknownRDR58
6,021 models trackedOpen

Research queue

arXiv

Preprints summarized through a developer-impact lens with source links and conservative claims.

  • Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesseshttp://arxiv.org/abs/2608.24876v1 - Zhaochen Yu, Yingcheng WuRDR86
  • From Seeing to Acting: Smart Glasses as First-Person Intelligence Platformshttp://arxiv.org/abs/2608.24877v1 - Jiangning Zhang, Haojun ChenRDR86
  • LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-traininghttp://arxiv.org/abs/2608.24845v1 - Andreas Hochlehnert, Marianna NezhurinaRDR84
  • Strictly Causal Streaming Video Anomaly Detection with a Theoretically-Grounded State-Space Corehttp://arxiv.org/abs/2608.24810v1 - Yogesh KumarRDR86
2,227 papers trackedOpen

Tool launches

Gated

Product and developer-tool launches with compliance-aware sourcing.

  • Z.ai confirms Ox Alpha is a new GLM-series model and will release its weightshacker-news-ai - Community discussionRDR64
  • RAG Is Simpler Than You Thinkhacker-news-ai - Community discussionRDR64
  • VMs won't contain cyber-capable agentshacker-news-ai - Community discussionRDR61
  • GLM-5.3-Flashhacker-news-ai - Community discussionRDR65
110 launches 7dOpen
Get the AI builder radar

Custom alerts and RSS for the segments you actually track.

Pick agents, AI coding, models, research, GitHub radar, providers, and a minimum score. Feeds use published, source-linked radar items only.

Topics
Choose segments and get a private RSS feed plus preference link.
How the radar works

Transparent scoring. No invented numbers.

53 active sources

Official blogs, GitHub, arXiv, Hugging Face, OpenRouter, RSS feeds, and public APIs with source links kept visible.

367 scans in 24h

Scan counts and signal totals come from the live radar, not from static homepage counters.

Radar score, 0-100

Scores combine reliability, freshness, novelty, technical importance, developer relevance, ecosystem signal, and confidence.

931 raw signals in 24h

Low-confidence items stay off public pages until their source support is strong enough.