Reddit r/artificial
The biggest AI bottleneck today with deployment layer is model iteration
One thing I've noticed while looking at production AI systems is that getting the first model deployed is rarely the hard part anymore. Most teams can build a AI apps like, support bot, document assistant, or agent workflow fairly quickly. The harder problem starts a few weeks later. Real users don't behave like benchmark datasets. They use internal terminology, ask incomplete questions, upload messy documents, and interact with systems in ways nobody anticipated during evaluation. As usage grows, you start seeing patterns: Certain questions consistently produce weak responses. New product terminology appears that wasn't in the original training data. Users find edge cases that never showed up during testing. The model performs well in some workflows and poorly in others. The problem is that most AI systems don't learn from any of this. Inference logs sit in one system. Training datasets live somewhere else. Fine-tuning pipelines live somewhere else. Evaluation is done using different tool. So every model improvement cycle becomes a project of its own. This is the biggest bottlenecks in production AI today. Not training but Model Iteration. Training is also a crucial part of it. Can you take production usage, identify failure patterns, turn them into datasets, improve the model, redeploy it, and repeat the process without rebuilding the entire workflow every time? The teams getting the most value from AI seem to be building feedback loops instead: production traffic → dataset curation → post-training → evaluation → redeployment Then repeating that cycle continuously. I recently tried the approach on one Insaurance chat usecase, and my pipeline kinda look like this: https://preview.redd.it/kdo9vytzfi6h1.png?width=1272&format=png&auto=webp&s=03d9799ace5a567eafd004a1d141084af6ee5afb I was looking at how platforms like Data Lab approach this problem recently, and the interesting part wasn't the fine-tuning itself. It was treating inference logs, datasets, post-training,
/u/codes_astro
2026-06-11 03:56
👁 5
查看原文 →
Reddit r/artificial
Claude Fable 5's security guardrails can be bypassed with a fake homework assignment
So Anthropic dropped Fable 5 yesterday with these hard blocks for anything security-related. Decided to poke at it. I asked it for help exploiting some vulns on a Metasploitable2 VM (it's a deliberately vulnerable training box, totally legal, it's mine). Fable 5 blocked it instantly and handed me off to Opus 4.8 as a fallback, which is apparently how it's designed. Opus 4.8 asked me to prove it was a legitimate request. So I spent 2 minutes writing a fake university course rubric — fake class, fake professor, fake Canvas deadline — and pasted it in. Opus 4.8 then gave me the full exploit walkthrough. Every command. Even offered to write my lab report for me. The guardrail works fine. The fallback is the hole. Anthropic essentially replaced "no" with "convince me" and the bar for convincing it is a Word doc you made up. Not reporting it because they don't pay for this. Sharing it here instead lol. https://preview.redd.it/o892vvv4fi6h1.png?width=1188&format=png&auto=webp&s=00e804d35e6cb4b672e036399c2c7e3ff7139f49 submitted by /u/dayumnn420 [link] [留言]
/u/dayumnn420
2026-06-11 03:51
👁 5
查看原文 →
Reddit r/artificial
Nobody needs AI to search the Internet, court says in ruling against Google
submitted by /u/Hot-Upstairs9603 [link] [留言]
/u/Hot-Upstairs9603
2026-06-11 03:51
👁 5
查看原文 →
Wired
A Meta Employee Who Just Lost Their Job Was Detained by Immigration Agents
Colleagues discussed the incident on internal message boards, according to documents seen by WIRED.
Lauren Goode, Paresh Dave, Louise Matsakis
2026-06-11 03:49
👁 11
查看原文 →
Reddit r/artificial
Dario Amodei — Policy on the AI Exponential
submitted by /u/Gloomy_Nebula_5138 [link] [留言]
/u/Gloomy_Nebula_5138
2026-06-11 03:36
👁 5
查看原文 →
Ars Technica
Google DeepMind releases DiffusionGemma, a model that runs local AI 4x faster
Diffusion AI is most common in image generation, but it can make text outputs much faster.
Ryan Whitwam
2026-06-11 03:29
👁 10
查看原文 →
HackerNews
Show HN: I am building a map of people who lived in the Roman Empire
Driving home from work one day, I wanted to know how many people we knew the names of who lived during the Roman era. Searching around, I found lists of Consuls and officials, but nothing that covered ordinary people or even most people like freedmen and slaves. So I ended up building a pipeline to process the more than 500k Latin inscriptions in the Epigraphic Database Clauss-Slaby https://edcs.hist.uzh.ch/en/ and extract the names of people (and attempt to cluster them, but this is a work in p
metiscus
2026-06-11 03:28
👁 2
查看原文 →
Reddit r/MachineLearning
Routing LLMs by task verifiability: a small experiment (n=120, 3 models) inspired by Karpathy's framework [D]
Full disclosure: this is directional, not a paper. n=120 tasks, one internal evaluator, not peer reviewed. I work at an LLM infrastructure company. This experiment was done on my own time and is not a company claim. Karpathy's framework classifies tasks by verifiability. Can output be mechanically checked? High verifiability tasks like code compilation and structured JSON extraction are safer because the verifier catches errors. Low verifiability tasks like creative writing are riskier. I wondered if high verifiability tasks are also easier in practice. Can a weaker model do them as well as a frontier model if the verifier catches mistakes? Setup was 120 tasks across four categories. Code unit tests, structured extraction, multi hop reasoning, creative summarization. Three models: Claude Sonnet 4.6, GPT 5.5, local Mistral 3 8B via vLLM 0.6.3. Pass rate for the first two, human rating 1 to 5 for the last two. Results were messy. Code unit tests: Sonnet 4.6 94%, GPT 5.5 91%, Mistral 3 8B 87%. With one retry Mistral 3 hit 95%. That surprised me. I expected the gap to be bigger. Structured extraction: Sonnet 4.6 97%, GPT 5.5 94%, Mistral 3 8B 89%. With retry 96%. Also closer than I expected. But here is where it got weird. Sonnet 4.6 initially scored worse than GPT 5.5 on structured extraction, which made no sense. Turns out our JSON schema had an ambiguous nested array that confused Claude's tool use parser. Fixing the schema brought Sonnet to 98%, but I kept the original numbers in the table because the mistake is part of the story. Your verifier is only as good as your schema. Multi hop reasoning: Sonnet 4.6 78%, GPT 5.5 71%, Mistral 3 8B 51%. Retry didn't help. The model would hallucinate reasoning paths consistently. This is where the capability gap was real. Creative summarization: Sonnet 4.6 4.2 out of 5, GPT 5.5 3.9 out of 5, Mistral 3 8B 3.1 out of 5. Expected. Interpretation: high verifiability tasks seem simpler in the sense that weaker model plus verifier ca
/u/DragonfruitAlone4497
2026-06-11 03:18
👁 8
查看原文 →
Reddit r/artificial
Thoughts on this Sam Altman quote?
“We see a future where intelligence is a utility, like electricity or water, and people buy it from us on a meter." What do you think this means in practice? Is this a reasonable vision for AI, or does it raise concerns about dependence on a few companies for access to intelligence ? submitted by /u/Choice-Scallion-3499 [link] [留言]
/u/Choice-Scallion-3499
2026-06-11 03:06
👁 5
查看原文 →
Engadget
Take a deep look at Halo: Campaign Evolved before it launches next month
This remake of the first Halo game's hits consoles on July 28.
staff@engadget.com (Lawrence Bonk)
2026-06-11 02:55
👁 13
查看原文 →
Dev.to
From an Empty Workspace to a Running Robot in One Prompt
The hard parts of robotics are supposed to be perception, planning, and control. So why does so much of the day go to everything that comes before them? The hidden setup tax in every robotics simulation project Ask anyone what's hard about robotics and you'll get the same list: perception, planning, control, navigation. The genuinely interesting problems. If you track where your hours actually go, though, a strange thing shows up. A big chunk of the day disappears before you reach any of that. You're not solving hard problems yet. You're just getting to the starting line: wiring up a workspace, writing description files, stitching together launch files, and coaxing a simulator into opening without errors. It's the unglamorous tax on every project, and most of us have quietly accepted it as the cost of doing business. Building a differential drive robot simulation in ROS 2 and Gazebo from scratch A diff drive base, a LiDAR, and Gazebo, set up from one prompt instead of an afternoon of boilerplate. A few days ago I wanted a simple mobile robot simulation. Nothing exotic: a differential drive base (two driven wheels, the classic mobile-robot setup), a LiDAR for sensing, running in Gazebo . This is the kind of thing that should be straightforward. In practice it's an afternoon of boilerplate before the robot so much as twitches. So instead of wiring it up by hand, I wanted to see how far Drift could get from a single prompt. To make it a fair test, I stripped the workspace down to nothing. No packages, no URDF, no launch files. A blank slate. Then I typed one line: "Create a mobile simulation from scratch." From XACRO to URDF: how the robot description gets generated in ROS 2 What the tool wrote first, and what XACRO and URDF actually do for your robot. It checked the workspace first: The opening move was sensible: it looked at the current directory to understand what it was working with. It generated a XACRO file for the robot's dimensions: XACRO is the macro-based for
Aditi Sharma
2026-06-11 02:46
👁 19
查看原文 →
HackerNews
Anthropic's Model Naming, Extrapolated
sammycdubs
2026-06-11 02:45
👁 8
查看原文 →
Engadget
You can personalize your Instagram algorithm now — unless you want to see more posts from accounts you follow
Instagram has expanded its algorithm personalization features to its main feed.
staff@engadget.com (Karissa Bell)
2026-06-11 02:44
👁 18
查看原文 →
Product Hunt
Bond
The AI to-do list that does itself Discussion | Link
2026-06-11 02:43
👁 5
查看原文 →
The Verge AI
Fable won’t answer basic biology questions
Anthropic just released Claude Fable 5, calling it the most powerful AI model it has ever made widely available and praising its skills in biology, among others. But the model won't answer basic biology questions - the kind you'd expect a high schooler to handle. Instead, it hands off the query to the former flagship […]
Robert Hart
2026-06-11 02:43
👁 11
查看原文 →
Dev.to
Most repos hit by the Shai-Hulud worm are still infected a week later, and the obvious fix punishes the victims.
This is a follow-up to my earlier posts, and it is more of an open question than an answer. I have the data, I have a way to act, and I am genuinely unsure that acting is the right call. I could use the community's help thinking it through. Last week a supply-chain worm got into my GitHub account and repositories. I got out, cleaned up the proper way, and wrote it up. Then I checked the public list of repositories hit by the same worm, to see how the cleanup was going across the ecosystem. Nearly a week later, most of them are still carrying the live payload. It is worse than a count When you look closely, a lot of the owners are clearly trying. But they are missing how this actually works, in two ways that matter: Deleting is not removing. They remove the malicious files with an ordinary commit. That takes the payload off the branch tip, but the commit that introduced it is still in history, and the blob is still recoverable by anyone who reverts or checks out the old commit. The only real removal is rewriting history (reset, not revert) and asking GitHub to purge the objects, because the fork network keeps them reachable by SHA. One branch is not all branches. They clean the branch they know about and never see the backdated copies the worm planted on other branches, which are still live. And the part that genuinely worries me: some of these owners are almost certainly opening the infected repository in VS Code or an AI assistant to fix it , which is exactly the trigger that runs the payload again. The act of trying to clean it can re-detonate it. So: a large number of repositories still carrying a live credential stealer, and a large number of owners and contributors who do not know they are still exposed. The dilemma Here is where I am stuck. There are two paths and I do not like either. Report them to GitHub. Their response is automated and blunt. The repo gets disabled, with no human in the loop, the same hands-off automation that locked me out of my own accou
Ionut-Cristian Florescu
2026-06-11 02:39
👁 11
查看原文 →
Dev.to
Debugging the Google Maps Duplicate Loading Bug in React
Originally published on clintech.me If you've integrated Google Maps into a React app and seen Autocomplete randomly stop working, Directions silently fail, or the API throw google is not defined on second render — you've hit the duplicate loading bug. Here's exactly what caused it in my case and how I fixed it. The setup that broke things While building delivery address flows at POLOM — a production e-commerce platform — I integrated Google Places Autocomplete across 20+ screens. I had the Maps JavaScript API loading in two places: A provider.tsx for global script loading across the app A useLoadGoogleMaps hook inside a shared component This caused race conditions. The Autocomplete and Directions APIs were initialising before the script fully resolved in some renders, silently failing in others. The failure wasn't consistent, which made it harder to catch. The fix Step 1 — Remove the global load Delete the script tag or next/script call in provider.tsx . There should be exactly one place the Maps API loads. Step 2 — Centralise in a hook Move all loading logic into a single useLoadGoogleMaps hook using dynamic loading. If you're on Next.js, next/script with strategy="afterInteractive" inside the hook is the right approach. Step 3 — Guard before initialising if ( ! window . google ?. maps ) return ; Check that the API is fully available before attempting to attach Autocomplete or Directions . Don't assume the script load event means every namespace is ready. Step 4 — Scope your ref correctly Bind the autocomplete instance to inputRef.current explicitly. If the component remounts, re-initialise the binding — don't assume the previous instance is still attached. The result One load, one source of truth, no race conditions. Autocomplete and Directions worked consistently across all 20+ screens without reinitialising on every render. Security — the step most developers skip Restrict your API key at the Google Cloud Console level: HTTP referrers: whitelist your domain onl
ATAYERO CLINTON
2026-06-11 02:38
👁 20
查看原文 →
Engadget
Ubisoft reportedly shuts down more studios and lays off staff in Barcelona and San Francisco
Ubisoft is reportedly trying to cut costs without leaving key franchises like Rainbow Six: Siege in the lurch with a new round of layoffs.
staff@engadget.com (Ian Carlos Campbell)
2026-06-11 02:37
👁 5
查看原文 →
HackerNews
Policy on the AI Exponential
yjp20
2026-06-11 02:36
👁 3
查看原文 →
Reddit r/webdev
How do you maintain project context when you hit limits and have to switch AI tools?
One of the most frustrating parts of using AI coding tools is what happens when you have to switch between them. Every time I move to a different tool, I end up re-explaining the project, architecture, requirements, previous decisions, and all the context I’ve already built up elsewhere. It feels inefficient having to repeatedly provide the same context every time I switch tools or start a new session. Is anyone else running into this? How are you handling it today? submitted by /u/Bar_David [link] [留言]
/u/Bar_David
2026-06-11 02:30
👁 5
查看原文 →