Dev.to
Datadog dashboards for prompt regression: the panels we actually keep
We wired our LLM eval suite into Datadog over about four months. Most of the panels we built got deleted. These are the five that stayed, and the metrics that feed them. TL;DR: We run an LLM-as-judge eval suite on every PR that touches a prompt, and we ship the results to Datadog as custom metrics. The dashboard started with fourteen panels. We kept five. The one that catches the most real regressions is per-criterion pass-rate split out by judge criterion, not the single rolled-up pass-rate number, because an aggregate of 91 percent hid the fact that one criterion had dropped from 0.95 to 0.62. Below are the metrics we emit, the Python that submits them, the monitor config we alert on, and the panels we tried and dropped. Some context on the setup so the rest makes sense. We are a Series-C dev-tool startup. We have a handful of prompts in production that do real work (classification, extraction, a summarization step in an agent loop). Each one has an eval set of tagged examples, somewhere between 80 and 400 per prompt. The judge is a separate model call that scores each output against a rubric. We run the suite in GitHub Actions. The eval job emits metrics to Datadog at the end of every run. Backend service health was already in Datadog, so putting eval data next to it meant one place to look during an incident instead of two. 1. Emit per-criterion pass-rate, not just the rolled-up number This is the one that earns its place. Our judge scores each output against multiple criteria. For the extraction prompt it is four: correct fields, no hallucinated fields, format valid, no refusal. Early on we only emitted one number, prompt_eval.pass_rate, the fraction of examples that passed every criterion. That number is fine for a smoke test and useless for debugging. The problem showed up on a prompt change that looked clean. Overall pass-rate went from 0.93 to 0.91. Two points. Nobody would block a PR on two points. But underneath, the "no hallucinated fields" criterion had
Ethan Walker
2026-06-09 02:18
👁 10
查看原文 →
Reddit r/artificial
Is AI Good or Bad? (Data Science Major)
I am a last-year data science major at university who initially joined because of AI's exciting potential across numerous industries. However, after learning about multiple companies backtracking on their AI use on their platforms and cutting back on their data center expansions, I can't help but think that something is very wrong behind closed doors. I came to understand that the demand for AI is slowly decreasing in some areas and increasing exponentially in others. To me, it seems every major industry "needs" AI to make life easier, yet is backtracking when it doesn't perform the way they want it to. My concerns revolve around how unpredictable AI's usage is. If I get involved in an industry that actively destroys land, water, and other resources, I would hope that the environmental costs will be outweighed by the benefits everyone sees from AI. However, with the economic trend of AI's value decreasing for companies that initially went all in on it, I can't help but feel like I'm actively destroying the planet. Does anyone have any suggestions or moral redemption for me? I want to jump ship before the big explosion, but I'll stay if there's great potential for growth with AI. submitted by /u/Emergency_Ad6929 [link] [留言]
/u/Emergency_Ad6929
2026-06-09 02:18
👁 6
查看原文 →
Dev.to
Anthropic: Claude Now Writes 80% of Its Own Code in 2026
80%. That is the share of code currently being merged into Anthropic's production systems that was written by Claude. Not code-reviewed. Not pair-programmed. Written. In February 2025, when Claude Code launched, that number was in the low single digits. Sixteen months later, the company decided that data point — and the trajectory behind it — was worth a public warning. On June 4, 2026, Anthropic published "When AI Builds Itself," a research paper co-authored by Marina Favaro, head of the Anthropic Institute, and Jack Clark, one of the company's co-founders. It was the first major publication from the Anthropic Institute since its founding in March 2026. The paper did two things simultaneously: disclosed internal productivity data that most AI companies keep private, and called for a global mechanism to slow or pause frontier AI development before the process becomes self-sustaining without meaningful human direction. The data came first. The policy recommendation followed from it. Here is what the numbers actually show and why every developer building on AI infrastructure today should read this carefully. The Productivity Curve Nobody Predicted Anthropic published a chart of engineering output per engineer, indexed to a baseline from 2021–2024. The curve is flat for four years. Then Claude Code shipped in February 2025. The multiplier progression from that point: 1.2x, 1.5x, 1.9x, 2.5x. By Q1 2026: 5.8x. By Q2 2026: 8x. The typical Anthropic engineer is now merging eight times as much code per day as they were in 2024. Not 8% more. Eight times more. That is not a productivity improvement — it is a different category of output from the same headcount. To understand what drives the number, you need to understand what Claude Code actually does inside Anthropic's engineering workflows. The tool was built for and by engineers working on frontier AI systems — which means the tasks it handles are not boilerplate CRUD endpoints. Claude is writing test harnesses for novel m
Anup Karanjkar
2026-06-09 02:17
👁 10
查看原文 →
HackerNews
Siri AI
0xedb
2026-06-09 02:17
👁 5
查看原文 →
Wired
Apple’s New Siri AI Is Ready to Get Personal
From a stand-alone app to a Google Gemini partnership, here’s everything you need to know from WWDC 2026 about Apple’s upcoming overhaul of Siri.
Reece Rogers
2026-06-09 02:17
👁 10
查看原文 →
Reddit r/MachineLearning
STOP racist posts about Chinese researchers [D]
Yes, I'm calling it out. It IS racism. As an active member of r/MachineLearning and a researcher who is ethnic Chinese, I am DISGUSTED by unfounded accusations against the group of researchers who constitute over half of the field. Such posts pop up every other week, grounded in conspiracy theories, and creating a sinophobia echo chamber. I understand the salty feeling when one's paper is rejected, no matter whether the paper actually deserves acceptance or not. Given the noise in conference organization and reviewing process, and a relatively junior body of participants, it is very likely that one finds a paper "worse than mine" slip into the conference, and there's a high chance that the paper has a Chinese author. That's simply because of the composition of the authors, and does not warrant accusations, aka witch hunts, towards certain ethnic groups. This sub is about an important scientific subject in the modern world. If anyone agrees with the logic "80% of the authors are Chinese, so my rejection is their fault.", they should seriously rethink their career plan since such thinking does not belong to serious scientists. We should be open to discussing the problems we have in the current conference organization and reviewing process, but racism should not have a foothold in our field. submitted by /u/AffectionateLife5693 [link] [留言]
/u/AffectionateLife5693
2026-06-09 02:11
👁 7
查看原文 →
Engadget
Apple Intelligence is coming to the Shortcuts app
Home and Shortcuts are not immune to Apple's AI upgrades.
staff@engadget.com (Anna Washenko)
2026-06-09 02:08
👁 5
查看原文 →
TechCrunch
Apple’s long-awaited AI Siri overhaul is finally here
The idea behind the new "Siri AI" is to turn the assistant from a voice controlled assistant into an AI companion that can do a lot more.
Aisha Malik
2026-06-09 01:56
👁 10
查看原文 →
The Verge AI
Apple is redesigning Screen Time and overhauling child controls
At WWDC 2026, Apple announced an overhaul to its Screen Time parental control features that aim to improve its safeguarding features to protect children who use iPhone, iPad, and Mac devices. Some of the new features coming with Apple's iOS 27, iPadOS 27, and macOS 27 updates include giving parents and guardians more control over […]
Jess Weatherbed
2026-06-09 01:55
👁 9
查看原文 →
Reddit r/artificial
Nvidia and SK Hynix Sign Multiyear AI Deal Ahead of Vera Rubin Launch
submitted by /u/andix3 [link] [留言]
/u/andix3
2026-06-09 01:54
👁 7
查看原文 →
Engadget
macOS Golden Gate puts Siri AI into Spotlight
macOS Golden Gate will add Siri AI into Spotlight.
staff@engadget.com (Devindra Hardawar)
2026-06-09 01:53
👁 6
查看原文 →
TechCrunch
Apple says it’s fixed the awful search function for emails, photos
Apple says a completely rebuilt Search function will competently find the emails, photos and other content you are searching for.
Julie Bort
2026-06-09 01:53
👁 8
查看原文 →
The Verge AI
Amazon is launching AI-generated custom merch
Amazon is expanding its print-on-demand features to AI-generated designs created using Alexa for Shopping for products like T-shirts, water bottles, and hoodies. Shoppers can use text prompts to generate images that are then printed on to blanks for sale on Amazon. They can then share the link to the design so other people can buy […]
Mia Sato
2026-06-09 01:52
👁 11
查看原文 →
Reddit r/webdev
CRM for appointments
A client of mine is starting a business and needs a CRM. Core functionality is to manage different professionals and assign customers to professionals to make appointments and manage agenda for each professional. I will assist this business on the software side and I am fully capable of developing a custom CRM on my own but it’s not convenient for time/money constraints. I have looked for open source CRMs I can extend so I don’t waste time on UI and common features. Are there products which can save development time? I wanted to evaluate Calendly and integrate it in the final solution, do you have any experience with this product? I would also evaluate existing products but open source I can fork would be best submitted by /u/Umberto_Fontanazza [link] [留言]
/u/Umberto_Fontanazza
2026-06-09 01:50
👁 7
查看原文 →
Reddit r/artificial
The AI productivity paradox that needs to be addressed rn
The conversation around AI coding is still stuck on velocity and its completely missing the real operational bottleneck -> DEBUGGING I use a combination of tools like GitHub Copilot, Cursor, and generic agentic code gen tools(whichever give me the most credits that week) , dropping a 300-line functional block from a natural language prompt takes about a minute. On paper, developer velocity should have been increased by 69 times. but i feel like the bottleneck hasn't disappeared; it just shifted down the pipeline. Like i traded manual work for incredibly frustrating debugging. LLM code looks fine on surface but like when u go through line to line, you feel like its built on sand i mean sure if it works it works but like one thing i struggle with is ghost features, like if i accidentally suggest a feature then the LLM is gonna shove it in my code, even if i say no later on. (if someone knows how to fix do dm) idk about ya'll but i'd much rather have a ai llm that takes like 1 hour to write 500 lines of code if that means i have to debug less. another thing how are you handling validation boundaries? are u using runtime timeout scripts or smth open source like gitagent? also this is gonna sound weird but i kinda have trust issues when a llm spits like 300-400 lines in under a minute (idk why) sorry for my bad english, im not a native speaker submitted by /u/SpicyTofu_29 [link] [留言]
/u/SpicyTofu_29
2026-06-09 01:44
👁 7
查看原文 →
Wired
The UK Is Betting on a Billion-Dollar AI Supercomputer to Kick Its Addiction to US Tech
The British government thinks a state-backed infrastructure initiative will help supercharge homegrown chip startups.
Joel Khalili
2026-06-09 01:44
👁 12
查看原文 →
Engadget
Apple reintroduces the AI-powered Siri it announced at WWDC 2024
At WWDC 2026, Apple resurrected its AI-powered voice assistant, now named Siri AI.
staff@engadget.com (Igor Bonifacic)
2026-06-09 01:41
👁 6
查看原文 →
HackerNews
HN seems dead compared to say 10-15 years ago
Dearth of original ideas, lots of pointless retro stuff like "I did x on mac os classic" lots of reinventing the wheel with LLMs and LLM cargo culting. What is the perpetual growth myth to do once physics is known and energy is constraining?
morpheos137
2026-06-09 01:37
👁 5
查看原文 →
The Verge AI
Apple announces Siri AI and its next generation of Apple Intelligence
Two years after first revealing its plans for Apple Intelligence and a smarter Siri that never fully materialized, at WWDC, Apple just revealed a new set of AI features and a smarter, more personalized Siri. Apple calls Siri AI an "entirely new version of Siri" and says it's both more conversational and more capable than […]
Dominic Preston
2026-06-09 01:34
👁 10
查看原文 →
Reddit r/programming
Making numpy-ts as fast as native
I started working on numpy-ts last year and began serious performance optimization in February. These are some of the challenges and lessons from this project. Some/all might be obvious - lmk what you think! Disclaimer: numpy-ts was written with some AI assistance. Please read my AI disclosure for more details. submitted by /u/dupontcyborg [link] [留言]
/u/dupontcyborg
2026-06-09 01:32
👁 7
查看原文 →