EveryDev.ai
Subscribe
Main Menu
  • Tools
  • Developers
  • Topics
  • Discussions
  • Communities
  • Users
  • News
  • Podcasts
  • Blogs
  • Builds
  • Contests
  • Compare
  • Arena
  • Polls
Rate toolsCreate
AI Tools by Topic
  • AI Coding Assistants
  • Agent Frameworks
  • MCP Servers
  • AI Prompt Tools
  • Vibe Coding Tools
  • AI Design Tools
  • AI Database Tools
  • AI Website Builders
  • AI Testing Tools
  • LLM Evaluations
Follow Us
  • X / Twitter
  • LinkedIn
  • Reddit
  • Discord
  • Threads
  • Bluesky
  • Mastodon
  • YouTube
  • GitHub
  • Instagram
Get Started
  • Users
  • Rate Tools
  • About
  • Editorial Standards
  • Corrections & Disclosures
  • Community Guidelines
  • Advertise
  • Contact Us
  • Newsletter
  • Submit a Tool
  • Start a Discussion
  • Write A Blog
  • Share A Build
  • Terms of Service
  • Privacy Policy
Explore with AI
  • ChatGPT
  • Gemini
  • Claude
  • Grok
  • Perplexity
Agent Experience
  • llms.txt
Theme
With AI, Everyone is a Dev. EveryDev.ai © 2026
    1. Home
    2. News
    3. Weekly AI Dev News Digest: September 26 - October 2, 2026
    Joe Seifi's avatar
    Joe Seifi
    October 3, 2026·Founder at EveryDev.ai
    Discuss (0)
    Weekly AI Dev News Digest: September 26 - October 2, 2026

    Issue #39 · Weekly Digest

    Weekly AI Dev News Digest: September 26 - October 2, 2026

    October 3, 2026

    Did you hear OpenAI released another model? It's called GPT-6.1 Sol.

    They're saying it's close to Astra for coding, but it costs about a fifth as much per token. Also, Anthropic released Sonnet 5.5 and it's supposed to get through the work with fewer steps. So in theory it costs less even though the token price hasn't changed. And at Rails World, people were arguing about who needs to read all this code.

    30%+

    faster Sonnet 5.5 output, Anthropic reports

    ·

    1/5

    Astra's standard token price for Sol

    ·

    1M

    Argon output-token limit

    ·

    64%

    lower median OpenSWE task cost in LangChain's experiment

    ·

    24

    Android vulnerabilities reported by GitHub's taskflows

    ·

    1M+

    databases launched weekly, Supabase reports

    In Focus

    New Models and the Cost of Coding

    The comparison with Astra comes from OpenAI's own tests. Sol matched it on DeepSWE, an evaluation of coding tasks in real repositories. So there's a reason to try the cheaper model on work that previously needed Astra. It still has to complete the change, though. If a cheaper run needs several retries, or somebody has to repair the result, the token price only tells part of the story. Sol is available in the API and on paid Work and Codex plans. ChatGPT access is still to come. (OpenAI)

    Sonnet 5.5 has a different way of bringing the cost down. Anthropic says it takes fewer steps and batches tool calls more effectively. There's also a detail in the evaluation notes: turning the effort all the way up sometimes made the result worse. The model brought in agents to review the code, and in some cases they ran out of time or made changes outside the request. The setting just below maximum got a better FrontierCode score. If the request was for one fix, extra changes can leave somebody with more to review. (Anthropic)

    In Focus

    Google’s Argon, and who gets to use it

    Google announced Gemini 4 Argon this week, but most developers can’t try it yet. It’s initially going to selected cybersecurity teams through Fairwind, a program that gives approved organizations access to models for finding and fixing security vulnerabilities. Google says Fairwind has more than 650 partners. Only some have Argon. Google’s program.

    People using it inside Google sound pleased. DeepMind’s Cristian Garcia said he’d used Argon for weeks and that it changed how he worked. But when someone hoped to get access soon, Garcia said he was waiting too—for his personal account.

    There’s skepticism, too. Ishu Agrawal questioned the enthusiasm on X, pointing to Google’s earlier Gemini 3.1 Pro benchmark chart. Google’s Logan Kilpatrick replied that most new Gemini versions now go through weeks of testing by thousands of Google software engineers.

    Artificial Analysis tested Argon outside Google. On its knowledge test, Argon at high reasoning answered 50% correctly, compared with 63% for Astra at maximum reasoning. Of the questions Argon didn’t get fully right, 15% received wrong answers. The rest received partial answers or weren’t attempted. Results, method.

    One Fairwind partner, Armadin, builds AI agents that try to break into customers’ systems with permission to find security weaknesses. It raised another $255.5 million on October 1. Its investors include Google Ventures, funded by Alphabet, and In-Q-Tel, an independent nonprofit investor that works with the CIA and other national-security agencies to find technology they can use. Those investments are public; their stakes aren’t disclosed in the sources we reviewed. Funding, GV, CIA.

    Armadin’s founder, Kevin Mandia, previously sold Mandiant to Google and joined Amazon’s board in September. CrowdStrike CEO George Kurtz sits on Armadin’s board. And Google owns Wiz, another featured Fairwind partner. Amazon, Armadin, Wiz acquisition.

    Another Fairwind partner, the Center for Internet Security, runs a security-support program for state and local governments. Some paying members could keep previously federally funded software that detects attacks on their computers at no extra charge through September 30. After that, CIS offers it as a paid service. CIS’s FAQ.

    Google says wider Argon access will start with paid API customers and Ultra subscribers. As of October 3, it hasn’t given a date. Rollout announcement.

    In Focus

    Choosing Which Model Does the Work

    A model doesn't always need to write a paragraph. Sometimes it just needs to pick a tool, sort a support ticket or decide whether a person should look at a request. Cloudflare's new Clef models are built for that kind of choice. They return answers with probabilities, without writing an explanation first. Clef and Clef-flash are available on Workers AI, and their weights are released under Apache 2.0. They're also compatible with Jev's API, so developers can compare them using the same interface. Customizing them with reinforcement learning starts with help from Cloudflare's engineers; self-service is planned later. (Cloudflare)

    LangChain tried routing in OpenSWE. It reads the first human message, decides how difficult the request looks, and picks a coding model. In the experiment, work cost less and pull requests were merged about as often as before. There's a limit to this version: it keeps that model for the whole conversation. A request can start as a small bug fix and become a much larger change. The router doesn't reassess it halfway through. That's something to test on your own requests, since the published result comes from OpenSWE's task mix. (LangChain)

    Sebastian Raschka's article gives some background on where these models come from. He goes through the history of classifiers and explains how Jev fits between a model trained for one specific decision and a general language model. It helps explain why a tool that doesn't write an answer can still be useful inside an agent. He also separates what's known about Jev from his guesses about how it's built. (Sebastian Raschka)

    Our Read

    A router needs to be measured against the work it actually receives. Include retries and escalations in the cost, and check what happens when the conversation changes direction.

    In Focus

    Agents That Keep Working While You're Away

    OpenAI's dots can keep working between conversations. They have their own cloud computers, and the company describes a developer's dot watching customer feedback, preparing fixes and bringing back pull requests with videos attached. That gives the person reviewing the change something to look at besides a written claim that it works. There are permission controls and action approvals, and access depends on the plan, market and administrator settings. A person coming back to several finished changes still needs to know what the agent tried and what needs a decision. (OpenAI)

    DevDay included more ways to build that kind of agent into other products. The Agents API gained computer use. Codex cloud environments became reusable, and plugin extensions can add interactive panels and file viewers. The Agents API was introduced earlier in September, so the news here is the added functionality. A persistent environment lets an agent continue its work. The panels and viewers give the user somewhere to inspect what it has produced without having to reconstruct everything from the conversation. (OpenAI DevDay)

    There's also the question of how to stop an agent doing something it shouldn't. NVIDIA's Open Agent Safety Platform puts controls outside the model. OpenShell sets boundaries around what the agent can do. The Sentry reference design watches activity from a BlueField-4 DPU, outside the host's trust domain. The agent's reasoning doesn't get to decide whether those controls apply. OpenShell software is available; Sentry is a hardware reference design. (NVIDIA)

    GitHub's security team published a more focused example of agents doing useful work. Kevin Stubbings describes audit taskflows aimed at Android entry points and particular types of vulnerability. The disclosed findings include location tracking through OsmAnd and account takeover in Wikipedia's Android app. Other researchers can run the open-source workflows on their own applications. They need a Copilot license, the runs consume premium requests, and a researcher still has to verify what the agent reports. The prompts and examples give that researcher a place to start. (GitHub Security Lab)

    In Focus

    Databases for Agent-Created Apps and the World Labs Deal

    If agents keep creating little apps and prototypes, each one may need somewhere to store data. Supabase says it's already launching more than a million databases a week. It announced that it's acquiring Turso, whose architecture can load small databases when they're needed and suspend them when they're idle. That avoids setting aside a machine for every small experiment. Supabase says it will keep building around Postgres, and Turso will continue its SQLite work. Glauber Costa will lead the agent infrastructure effort with the joining team. (Supabase)

    Supabase Select also included ways for a coding agent to handle more of the backend from the repository. Schema files and configuration can live alongside the code. Compute services can run for longer periods, and an app can have an authenticated MCP server for its users' agents. That last part needs a little explanation: Supabase's developer MCP server helps build an app. The app's own server lets a user's agent do things inside it, under that user's permissions. Local development without Docker is also available in alpha, off by default. (Supabase Select)

    World Labs signed an agreement to join AMD. Fei-Fei Li is set to become executive vice president and chief scientist, reporting to Lisa Su, while Justin Johnson and Ben Mildenhall continue leading model research. The team works on spatial intelligence: models concerned with the spatial and physical world. It has already been working with AMD on training and inference. Joining the company would bring that model research closer to the hardware and software used to run it. The deal is expected to close by year-end, subject to approvals, so it hasn't been completed yet. (World Labs)

    Why This Matters

    Creating a database for an experiment is one thing. Somebody still needs to know who owns it, when to delete it and what happens if the experiment becomes a real application.

    In Focus

    The Argument at Rails World

    At Rails World, the disagreement was about how much code people should hand over to agents, and how much they still need to understand. David Heinemeier Hansson described using generated Rust for HEY infrastructure and moving away from writing code by hand. Joey deVilla wrote a sharply critical response on September 27. The opening video is from before this week; the response and closing talk are new. (Global Nerdy)

    Aaron Patterson's September 27 closing keynote takes the other side. He talks about the work that goes into Ruby's runtime and compilers, and makes a distinction between a compiler and an AI agent. A compiler can change how a program runs internally while preserving its observable behavior. An English request to an agent doesn't come with that same promise. Patterson argues for continuing to read the code and understand it. He also jokes about handing colleagues large generated changes, which gets at who ends up doing the review after somebody else has finished generating. (Rails World)

    Cal Newport asks about responsibility at the labs themselves. In a September 28 essay, he argues that specific risky agent experiments and the safety procedures around them should be investigated. He wants scrutiny of what the labs are actually doing, rather than treating AI's future as inevitable. It's his argument for investigation, not a settled account of the labs' conduct. His question is who gets to examine those experiments and the decisions behind them. (Cal Newport)

    Our Read

    Watch the Rails talks together. They disagree about how much to hand over, but someone still needs to understand what the code does well enough to review it and fix it later.

    Signals

    Signals from the Edges

    Pi has a stable release and an experimental companion

    Earendil released Pi 1.0 with deferred tool loading, virtual-model extensions and ways to change prompts and tools while keeping track of the conversation. Pi Durable is a separate package for longer-running applications, and it's still experimental. Both are MIT licensed. Pi can be the terminal agent; Durable gives developers another option for building an application around it.

    Earendil→

    AstaBrief writes reports from retrieved research

    Ai2 released AstaBrief 8B and its training data, plus an example for working with local PDFs. It takes a question and excerpts from papers and writes a cited report. Something else still has to find those papers. Ai2 also says most of the training and evaluation happened in 2025, and it hasn't rerun the full comparison against today's frontier models.

    Ai2→

    DeepMind is putting watermarks into proteins

    SynthID Bio marks AI-designed biological sequences and structures while preserving their function in laboratory tests. The idea is to follow the provenance of a design into the physical protein, beyond the digital file. It's a proof of concept, with research materials released for others to use. Making the marks hold up against deliberate tampering is still an open problem.

    Google DeepMind→

    Sam Witteveen demonstrates decisions inside an agent

    His September 29 video, Using Jev in Your Agent Harness, covers routing, risk checks, tool selection and retrieval ranking. The published chapters include skill selection and RAG reranking demos, plus compound questions, prompt injection and situations where Jev doesn't fit. There's enough detail to choose a particular part of an agent to experiment with.

    Sam Witteveen→

    A move between coding-agent companies led to a public dispute

    Chris Degnan's move to Cognition prompted accusations from Factory's Matan Grinberg about advisory work overlapping with recruitment discussions. Degnan and Scott Wu disputed the account. Dr. Web's October 2 report follows both sides, along with opposing responses from Vinod Khosla and Keith Rabois, whose firm backs both companies. The sequence of events and the conduct involved remain disputed.

    Dr. Web→

    Fireship covers a preview of information-flow controls

    Its September 30 video looks at OpenAPPA. It checks tool calls against the sensitivity and trust of data an agent has read. The check is deterministic, so the same events produce the same policy decision. The surrounding model can still behave differently, and the project is a preview. Those details matter when assessing the claim that prompt injection has been solved.

    Fireship→

    Looking Ahead

    What to Watch

    1. 1

      When more people can try Argon

      Google's engineering examples give developers something to compare against once they have access. The results to watch are completed changes, the review they need and what it costs to get there.

    2. 2

      When a routed task gets bigger

      A model chosen for the first message might not suit the request ten messages later. Watch how routers decide when to reassess or hand over to another model.

    3. 3

      What happens to the temporary databases

      Cheap creation is useful. Ownership, cleanup and recovery will matter as more experiments become applications people depend on.

    4. 4

      What an agent brings back for review

      A recording, a pull request and a clear account of the decisions can help someone check the work. Watch whether that part improves alongside the agent's ability to keep working.

    5. 5

      How teams handle generated changes

      The Rails debate has a practical follow-up: how much code people are expected to review, how well they understand it and whether they can maintain it afterward.

    Writing the code is getting cheaper. Understanding what changed still takes time, especially when it lands in somebody else's pull request queue.


    About the Author

    Joe Seifi's avatar
    Joe Seifi

    Founder at EveryDev.ai

    Apple, Disney, Adobe, Eventbrite, Zillow, Affirm. I've shipped frontend at all of them. Now I build and write about AI dev tools: what works, what's hype, and what's worth your time.

    Comments

    No comments yet

    Be the first to share your thoughts