Read 355+ AI Discussions
Recent activity · New posts and replies
Growing faster than ever now
The agentic Code Review space is very hot right now Feels like the river of AI-written code is getting wider, and review is the next narrow stretch. Copilot’s Dynamic Workflows and Kiro’s workflows let you put review steps into the process. I’d want a failed check to actually stop the change, otherw…
Why a clean average can hide a failed AI answer
A batch average does not tell a reviewer which answer must not ship. Separate sample severity from an overall score. Consider two fictional outputs. One has harmless extra punctuation. The other gives a 30-day refund window when the supplied policy says 14 days. An average can make both look like sm…
I'd start by making each relation instance its own node, with named links to its arguments. So f : A → B becomes a relation node connecting f, A and B, with roles like “function,” “domain” and “codomain.” You can represe…
Can you read the actual comments in ReviewBench?
Can you read the actual review comments in ReviewBench? The false positives are what I'm curious about, cause some are obvious nonsense you can skip, others sound convincing and now you're digging through the code looking for a bug that isn't there...
Might give Beam a go
Might give Beam a go once access opens up. Already getting carried away thinking about what to build with it, and they haven't even told us what it costs yet 😅
Can an agent figure out what went wrong on a real phone?
Oh, Android CLI connects the remote phone to ADB automatically now 👀 curious how an agent handles a bug that only shows up on one phone. Can it figure out what happened from the screenshots, or does it get stuck?