Reddit r/MachineLearning
Introducing Papers Without Code [P]
Hi, Niels here from the open-source team at Hugging Face. I've recently relaunched paperswithcode.co as a source for finding the state of the art (SOTA) across various AI domains, from 3D generation to AI agents. This is done by automatically parsing research papers published on arXiv/Hugging Face, enabling leaderboards to be created. See BrowseComp below as an example (a scatter plot and a table are available for each benchmark). - Scatter plot (you can hover over the dots to see the models): https://preview.redd.it/9rz2r3ffcf6h1.png?width=2880&format=png&auto=webp&s=b3f8e7a870802f6ef8227ecc0619e9e1057554b0 - Table: https://preview.redd.it/qoqriddw5f6h1.png?width=2862&format=png&auto=webp&s=a0034574f693847537037013672fb61daf27b16e As you can see, I've added support for viewing evals for closed-source models, too, given that many benchmarks are nowadays dominated by them, like GPT-5.5 and Mythos 5. You can always disable viewing closed-source evals with a toggle or in your PwC settings: https://preview.redd.it/p3k6jt6q6f6h1.png?width=1582&format=png&auto=webp&s=40149e51d6b326a77e53e33baf70d9850b3de365 When you turn them off, here's what the open model leaderboard looks like: https://preview.redd.it/tg42sin36f6h1.png?width=2838&format=png&auto=webp&s=1330a117ae9b4e0ce6d459493ae9e8f64107310a Closed-source papers are treated as regular "papers", although they can be any source, like a blog post (given that PwC supports submitting any source beyond arXiv). See the GPT-5.5 or Mythos 5 papers as examples, with their evals at the bottom. Notice the "closed" tag on their evals. Hence, you could jokingly call these "papers without code". Let me know what you think of this, and whether anything needs to be changed or added! Kind regards, Niels submitted by /u/NielsRogge [link] [留言]
/u/NielsRogge
2026-06-10 16:58
👁 7
查看原文 →
Product Hunt
SlimSnap
Your AI doesn't know which button you mean Discussion | Link
Alexander Bickov
2026-06-10 16:43
👁 3
查看原文 →
The Verge AI
WhatsApp ordered to host rival AI assistants for free
Meta has been ordered by the European Commission to restore free WhatsApp access for chatbots made by rival AI providers while the regulator finishes its antitrust investigation. The rare interim measure announced on Tuesday was deemed necessary "to prevent serious and irreparable damage to competition" in the general-purpose AI assistant market. This is only the […]
Jess Weatherbed
2026-06-10 16:40
👁 10
查看原文 →
HackerNews
AWS Bedrock to require sharing data with Anthropic for Mythos and future models
> For Fable 5, Mythos 5, and future models on Bedrock with similar or higher capability levels, Anthropic will require 30-day retention for all traffic on Mythos-class models. Retaining data for a limited period allows Anthropic to detect patterns of misuse that are not visible from a single exchange. Once you opt into data retention, your data will leave AWS’s data and security boundary. From the announcement here: https://aws.amazon.com/blogs/aws/anthropic-claude-fable-5-on-aws-mythos-class-ca
TomAnthony
2026-06-10 16:21
👁 3
查看原文 →
HackerNews
Germany's €100B bid to make the trains run on time
JumpCrisscross
2026-06-10 16:04
👁 3
查看原文 →
Product Hunt
Easybilling
AI-native billing & payments for usage-based AI products Discussion | Link
Siqi Shi
2026-06-10 15:48
👁 3
查看原文 →
Reddit r/webdev
Chrome 149 finally lets you turn off its local AI model. That should be the default
Google pushed a 4GB local AI model to Chrome through silent updates and did not provide a disable switch until version 149. Users had to delete the file manually and it would be re-downloaded on restart. The reason this matters is not the storage. It is the consent. An AI model running in my browser is a category different from a calculator widget. It sends data to an inference engine, consumes power, generates heat, and runs code. Not having a clear off switch is not an oversight. It is a product philosophy about whether the user is in control. I do not think local AI is inherently bad. For real-time search suggestions or on-device content filtering it is useful. But the deployment model matters. If I install something, I should know what it does and how to turn it off. The update that installed the model was silent and the documentation was buried. The switch to disable it only appeared after sustained user complaints. The lesson is that capability is not what builds trust. The ability to turn it off is. submitted by /u/Fantastic-Place5501 [link] [留言]
/u/Fantastic-Place5501
2026-06-10 15:48
👁 5
查看原文 →
TechCrunch
Meta signs first AI data center deal in India with Reliance
The 168-megawatt facility will support Meta's global AI computing needs and can be expanded over time.
Jagmeet Singh
2026-06-10 15:05
👁 8
查看原文 →
The Verge AI
Logitech’s new Mobi Fold squeezes a lot of functionality into a tiny folding mouse
Logitech finally announced its new ultraportable travel mouse following leaked marketing images that spoiled the surprise last month. As the name implies, the Mobi Fold is a compact mouse that can fold in half using a hinge that can pivot about 130 degrees. At $79.99 in graphite, off-white, lilac, and sand color options, the Mobi […]
Andrew Liszewski
2026-06-10 15:01
👁 8
查看原文 →
Dev.to
Cache Deep Dive IV — TLB, Huge Pages, and Memory-Level Parallelism
Earlier parts examined the performance characteristics of sequential and random access under single-threaded execution, and noted in passing the destructive effect of random access on the TLB. This part devotes full attention to the TLB: what it is, why a TLB miss is more severe than a cache miss, why a page table walk constitutes one of the longest dependency chains a CPU can encounter, how huge pages fundamentally alter TLB reach, and where memory-level parallelism falters in the face of TLB misses. Page Boundaries: Where the Prefetcher Halts Part III, in its discussion of prefetchers, noted a hard constraint: a prefetcher must not cross page boundaries on its own authority. The operating system manages virtual memory in units of pages (typically 4 KB, i.e., 64 cache lines). When a program reaches the end of one page and is about to step into the next, the prefetcher cannot proceed. The reason is that the next page may not reside in physical memory (it may have been swapped out to disk), or it may be an entirely invalid virtual address — if the prefetcher were to speculatively initiate an access to the next page, it would trigger a page fault: the OS would have to suspend the process and swap the page in from disk; in the case of an invalid address, the OS would terminate the process outright. From a security standpoint, the prefetcher neither can nor is permitted to autonomously cross page boundaries without TLB approval. Hence a performance brake appears every 4 KB — even when traversing an array sequentially, after every 64 cache line accesses the prefetch pipeline must pause and await confirmation of an address translation. This is not to say that modern CPU prefetchers are completely unable to cross pages. Intel's Next Page Prefetcher and AMD's equivalent mechanism can consult the TLB when approaching a page boundary — if the address mapping for the next page is already registered in the TLB, the prefetcher receives clearance to continue prefetching across th
Tyler Tan
2026-06-10 14:55
👁 14
查看原文 →
Dev.to
How to Transcribe a YouTube Video (Free, in Under a Minute)
Building a "paste a YouTube link, get a transcript" feature sounds trivial until you deploy it to a server. The moment your request comes from a datacenter IP instead of a residential one, YouTube responds with LOGIN_REQUIRED or quietly serves nothing. Here's how VidTranscriber handles it. The problem There are two ways to get text from a YouTube video: Existing captions — if the uploader (or YouTube's auto-caption) provides them, you can fetch the caption track directly. Fast, free, no transcription needed. Transcribe the audio — pull the audio stream and run it through a speech-to-text model (Whisper-family). Works for any video, but costs compute. Both start with talking to YouTube from your server — and that's where it breaks. YouTube aggressively gates datacenter traffic: the watch page and InnerTube API return LOGIN_REQUIRED , and naive audio fetching gets reCAPTCHA'd. The approach The fix is to separate where the request originates from where the work happens : A Cloudflare Worker handles the user request and orchestration. Caption/audio fetching is routed through a path whose egress isn't treated as a bot — so the LOGIN_REQUIRED wall doesn't trigger. Captions, when available, become the primary path (no transcription cost). Only when there are no usable captions do we fall back to downloading audio and running Whisper. Long jobs go onto a queue (Cloudflare Queues) so the request returns immediately and the transcript streams in as it completes. Why captions-first matters Most "transcript generator" traffic is for videos that already have captions — talks, tutorials, news. Serving those from the caption track is instant and free, which means the expensive Whisper path is reserved for the minority of videos that actually need it. That's the difference between a tool that's cheap to run and one that isn't. What's still hard IP reputation drifts — what works today can get throttled tomorrow, so the extraction path needs monitoring and fallbacks. Caption quality
Terry Shine
2026-06-10 14:55
👁 22
查看原文 →
Dev.to
Headless CMS Security: Why Decoupled Is Safer
📝 Originally published on unfoldcms.com — reposted here for the DEV community. (I work on UnfoldCMS.) A coupled CMS puts the admin login on the same hostname visitors reach. A headless CMS puts it on a different hostname behind auth. That single architectural difference is why headless CMS security is meaningfully better than traditional coupled-CMS security on most real-world dimensions — and it's also why the comparison gets oversimplified into "headless is more secure" when the truth is more interesting. This post is the architectural take on headless CMS security : why decoupled is safer on most dimensions, where it can be less safe if you don't handle API hygiene properly, and what the honest comparison looks like in 2026. TL;DR : headless wins on attack-surface reduction (admin off the public hostname, smaller plugin attack surface, API-first auth model) but loses on dimensions teams typically don't think about (exposed APIs without rate limits, JWT misuse, secrets in frontend code, draft preview tokens leaking). A well-built headless CMS is meaningfully more secure than a typical WordPress site; a poorly-configured headless CMS can be worse than a maintained WordPress site. The architecture biases toward safer; the implementation determines actual outcomes. The audience: technical decision-makers and security-conscious teams comparing CMS architectures with security as a deciding factor. If you're earlier in the architectural decision, see headless CMS vs traditional CMS: key differences . For the WordPress-specific security picture this post compares against, WordPress security problems in 2026 . The Attack Surface Difference The single biggest architectural difference between coupled and headless CMS security is where the admin lives . A traditional WordPress site puts the admin login at yourdomain.com/wp-admin . The same hostname your visitors reach. The same SSL cert. The same Cloudflare config. Every brute-force attempt, every credential-stuffing bot, ev
hamed pakdaman
2026-06-10 14:54
👁 13
查看原文 →
Dev.to
WordPress Market Share Declining (2026 Data)
📝 Originally published on unfoldcms.com — reposted here for the DEV community. (I work on UnfoldCMS.) WordPress's market share among CMS-using websites dropped from 65.2% in 2023 to 60.2% by Q1 2026, contracting -2.9% year-over-year for the first time in over a decade. The platform that built the modern web is losing share to Wix, Squarespace, Shopify, and an emerging headless category — and the trend is steeper among newly-built sites than the headline number suggests. This post is the data-driven take on WordPress market share declining in 2026 — where the numbers actually come from, where the lost share is going, what age-cohort analysis reveals about the trajectory, and what the developer-signal data (Stack Overflow, GitHub) shows about who's still building on WordPress versus moving on. TL;DR : WordPress isn't collapsing — 60% market share is still dominant — but the lead has shrunk for the first time in 10 years, the cohort of new sites is breaking away faster than the overall number suggests, and the developer mindshare is leaving even faster than market share. The structural pressures (security, performance, plugin tax, modern stack expectations) all point the same direction. The audience: developers, agencies, and CTOs trying to read the WordPress market trend before committing to a multi-year platform decision. If you've been told "WordPress is sinking" or "WordPress is fine, market share gossip is overstated," this post puts numbers behind the actual movement. For the broader context on why WordPress is losing share, see why developers are leaving WordPress: 7 pain points and WordPress vs modern CMS: honest feature comparison . The Headline Numbers Three primary sources track CMS market share. They don't agree exactly, but they all show the same direction: W3Techs (the most-cited dataset, scans the top 10M websites): Metric 2023 Q1 2024 Q1 2025 Q1 2026 Change WordPress share among CMS-using sites 65.2% 64.1% 62.4% 60.2% -5 points in 3 years WordPress shar
hamed pakdaman
2026-06-10 14:53
👁 15
查看原文 →
Dev.to
🤖 Your AI Agent Is Failing in Prod — You Just Don't Know It Yet
The demo is impressive. ✅ The demo works in your environment, with your data, with you watching. ✅ Production? Silent failures. Cost overruns. Wrong tool calls. Stuck loops. No fallback. ❌ Agents in 2026: The Real Problem Here is the thing most people are not talking about when they ship AI agents: A demo agent and a production agent are completely different things. A demo is: "watch this work once." A production agent is: "what happens when it is wrong, stuck, expensive, over-permissioned, or called 10,000 times by real users?" That second question is what separates a cool technical proof-of-concept from something a business can actually rely on. Demos are not systems. 1️⃣ The 7 Things That Break in Prod In every agent hardening sprint I run, the same failures show up: Failure Mode What It Costs No logging You have no idea what the agent did or why No eval set You cannot measure quality or catch regressions Unlimited tool access Agent calls tools it should never touch No retry logic Transient failures become permanent failures No memory rules Context leaks between sessions or inflates cost No fallback path Agent loops or crashes instead of escalating No cost checks 1 misconfigured prompt → $400 API bill overnight If your agent is in production with 3 or more of those missing — you are one bad prompt away from a very expensive incident. 2️⃣ The Production Hardening Checklist Before you call an agent production-ready, run through this: Eval set exists — at least 20 test cases covering happy path + edge cases Structured logging — every tool call, every input, every output, every error — logged and searchable Retry logic — transient API failures handled gracefully, not crashed Tool limits — agent cannot call tools outside its defined scope Memory rules — what carries over between sessions, what gets cleared, how context is compressed Fallback paths — when the agent gets stuck or uncertain, it has an exit: escalate to human, return partial result, surface an error Cost
CyprianTinasheAarons
2026-06-10 14:52
👁 12
查看原文 →
Dev.to
⚡ Proof Compounds. Claims Decay. — Why Delivery Is Your Next Marketing Asset
Here is the move most technical service providers miss: Every project you deliver quietly dies inside a private folder. Every project you deliver with receipts becomes a trust asset that sells the next sprint without you lifting a finger. The Insight Almost No One Acts On Delivery is not the end of marketing. Delivery is where the next marketing asset is born. The before/after screenshot. The launch-readiness report excerpt. The workflow map. The metric improvement. The buyer quote. All of that is proof. And proof is the compound interest of service work. Claims decay. Proof compounds. 1️⃣ What Proof Actually Looks Like This is the proof asset menu. Every sprint should produce at least 1 item from this list: Before/after screenshot — the most shareable format Launch-readiness report excerpt — shows rigor and standard Workflow map — visual, specific, credibility-dense Dashboard screenshot — metrics that moved Test checklist — shows what was verified, not just what was built Client quote — even 1 sentence is worth 1,000 words of claims Metric improvement — "response time dropped from 24 hours to 4 minutes" Public teardown — anonymous version of the diagnosis Case study — structured story: context → pain → fix → result One-minute walkthrough video — screen-recorded, narrated, personal You do not need all of them. You need 1 per sprint. 2️⃣ The Case Study Structure That Sells A case study is not a trophy. It is a reusable trust asset. Use this structure every time: 1️⃣ Context — who had the problem? (anonymized if needed) 2️⃣ Pain — what was it costing them? 3️⃣ Hidden cause — what was really broken underneath? 4️⃣ Fix — what did you change, specifically? 5️⃣ Result — what improved? With a number. 6️⃣ Proof — what artifact backs it up? 7️⃣ Lesson — what should similar buyers do next? That is 7 steps. The whole thing can fit in a LinkedIn post or a page section. And here is the thing most people are not talking about: a case study with a specific number outperforms 10 po
CyprianTinasheAarons
2026-06-10 14:52
👁 10
查看原文 →
Dev.to
🔥 The Sales Call Is Not a Performance — It's a Diagnosis
I have watched founders lose sales calls they should have won. Not because they lacked skill. Not because the offer was wrong. Because they walked in to prove they were smart — instead of finding out whether the pain was real. Sales Is Diagnosis Plus Decision The call is not there for you to pitch. The call is there to find out: Is the pain real? Does the buyer have urgency? Does the budget exist? Can a fixed-scope sprint create a clear win? That is it. Four questions. Everything else follows from those. Sales is not pressure. Sales is diagnosis plus decision. 1️⃣ The Call Structure That Works Frame the call in the first 60 seconds: "I'll understand the current state, ask what is costing you, then tell you whether a sprint makes sense. If it doesn't, I'll say so." That sentence does 3 things: Sets expectations — no pressure, no hard close Signals competence — you have done this before Removes the buyer's guard — they can be honest about what is broken Then run this flow: 1️⃣ Current state — what exists now? 2️⃣ Pain — what is broken or slow? 3️⃣ Cost — what does it cost in time, money, trust, or delay? 4️⃣ Urgency — why now? 5️⃣ Decision — who approves? 6️⃣ Success — what would make this worth paying for? 7️⃣ Close — recommend the sprint or walk away 2️⃣ The Questions That Reveal Money These are the 6 questions I use to find whether a sprint is worth recommending: "What happens if this stays broken for another 30 days?" — reveals urgency and cost "What have you already tried?" — reveals how serious they are "Where does the current process lose leads, users, time, or trust?" — reveals the money leak "Who feels this pain most inside the business?" — reveals whether the buyer is also the decision-maker "What would make this an obvious win?" — reveals success criteria before you price "If we fixed only one thing first, what would matter most?" — reveals scope Listen for the answer with the money in it. That is the thing you fix. That is what you price. That is the sprin
CyprianTinasheAarons
2026-06-10 14:52
👁 7
查看原文 →
Dev.to
🚀 Outbound Without Begging — The Contextual Outreach System That Works
Cold outreach fails when it feels like a stranger asking for your time. It works when it feels like a useful operator noticed a real problem and offered a small, low-risk next step. The difference is almost always structure. Not charisma. Not volume. Structure. The Conversation I Keep Having I see founders send 100 cold DMs with zero replies. Then I read the messages. "Hi [Name], I help businesses with AI automation. Would love to connect and explore synergies." That message fails on every line: No context — why this person, why now? No observation — what did you actually notice? No value — what is the useful thing? No risk — what is the low-friction next step? No diagnosis — you sound like you want something, not like you see something Here is the thing most people are not talking about: the DM that gets a reply is the one that feels like it was written about the specific person reading it. 1️⃣ The Outbound Formula Every message that works uses this structure: Part Purpose Context Why this person, why now? Observation What did you actually notice about their product/page/content? Risk or opportunity What might be costing them that they haven't seen? Useful next step Checklist, teardown, quick audit — something useful Light CTA Easy to answer — not a marriage proposal The goal is not to sell on the first message. The goal is to be useful enough that they want the next thing you send. 2️⃣ Three Messages That Actually Work AI app launch opener: "Saw your launch. Before adding more features, I'd check the hidden trust risks: auth, payments, logging, analytics, and onboarding. Want the launch-readiness checklist?" GTM system opener: "Your product looks useful, but the path from attention to booked calls feels thin. I can map the missing GTM system — want a quick look?" Workflow automation opener: "There's probably money leaking between first enquiry and follow-up. I can show you the 7-day missed-lead recovery workflow if helpful." Notice what they all share: Specific —
CyprianTinasheAarons
2026-06-10 14:51
👁 5
查看原文 →
Dev.to
🧠 The Million-Dollar Math Is Boring — And That's the Point
A million dollars is emotional as a dream. As math, it is boring. And that is exactly why most people never get close. Break It Down Here is the thing: $1M/year is not one big bet. It is a machine. And machines are built from boring, repeatable components. 20 clients at $50,000? That is $1M. 100 clients at $10,000? That is $1M. 12 retainers at $4,000/month? That is $576k — plus 4 sprints at $10,000 each gets you to $616k. The question is not whether the number is possible. The question is which machine can realistically produce it — from where you actually stand today. 1️⃣ The Practical Ladder Here is how the staged path actually works for an AI service business: Stage What You Are Doing Why It Matters Stage 1 Sell fixed-scope sprints Creates cash and proof Stage 2 Turn repeated sprint work into templates, SOPs, automations Reduces delivery time, increases margin Stage 3 Sell retainers around highest-demand system Predictable monthly cash Stage 4 Productize repeated workflow into software or toolkit Scalable without more hours Stage 5 Scale the thing the market already proved it wants Compound the machine Notice what is missing from Stage 1. There is no SaaS. No product. No cold paid traffic. No team. Just skill, packaged cleanly, sold to people with money and a painful problem. That is the fastest path — not the most glamorous one. 2️⃣ The Proof-of-Force Line The first mission is not $1M. The first mission is $10k/month — reliably, from sprint work. Here is what that actually looks like: 2 × $1,500 teardown/audit packages = $3,000 2 × $3,500 implementation sprints = $7,000 2 × $5,000 launch/GTM sprints = $10,000 3 × $2,000 retainers = $6,000/month That is not the finish line. It is the proof-of-force line. It proves the machine works. It funds the next iteration. It creates the case studies that make the next sprint easier to sell. Then you go from $10k/month to $25k. Then $50k. Then you make the productization decision from a position of demand — not hope. 3️⃣ The
CyprianTinasheAarons
2026-06-10 14:51
👁 6
查看原文 →
Dev.to
📊 Distribution Is the Moat — And Most Technical Founders Have None
Products are easier to build. Workflows are easier to automate. Content is easier to generate. But trust is not easier. Attention is not easier. Buyer memory is not easier. The Hard Truth Here is the thing most people are not talking about in 2026: The bottleneck is no longer the product. The bottleneck is whether the right buyer has seen your diagnosis 3 times in 2 weeks. Because that is how trust is built. Not with one perfect post. With repeated, useful presence in the right feed. Distribution is the moat. 1️⃣ Why "Staying Active" Is the Wrong Goal Most founders post to stay active. That is not a content strategy. That is anxiety dressed up as marketing. Every post should do one of 3 things: Make the buyer understand a pain they already have Make the buyer trust your diagnosis of that pain Move the buyer closer to a conversation A post about your tech stack? Probably none of those. A post that says "Your AI app is not launch-ready until auth, payments, logging, and rollback are boring" — that does all 3. 2️⃣ The Five Content Pillars That Build Pipeline Here is the system I use. 5 pillars. Everything maps to one of them: Pillar What It Signals Launch risk Why AI-built products break before production GTM systems How founders turn expertise into pipeline Workflow automation How businesses leak time and revenue Proof and case studies What changed before/after — with receipts Founder operating lessons The discipline behind building for money Every post I write maps to one of these. Not because it is tidy. Because each pillar speaks directly to a buyer who has a specific pain — and positions me as the operator who sees it clearly. 3️⃣ The Daily Format That Creates Pipeline This is the actual weekly posting structure that works: Monday — mistake post: a painful thing technical founders do wrong Tuesday — teardown post: a real example dissected publicly Wednesday — checklist: the 10-item audit your buyer needs Thursday — before/after: what changed after a sprint, with s
CyprianTinasheAarons
2026-06-10 14:51
👁 7
查看原文 →
Dev.to
🎯 "I Build AI Automations" Is Killing Your Close Rate
The conversation I keep having with AI founders goes like this: "I've sent 50 DMs. No one is biting." Then I look at the offer. "I build AI automations for businesses." There is the problem. Bad, Better, Best — The Offer Anatomy Breakdown Most technical people sell their skill. Buyers do not buy skill. They buy a removed headache. Let me break down the bad/better/best framework I use for every offer I build: Bad: I build AI automations for businesses. Better: I help service businesses automate lead follow-up so no enquiry gets ignored. Best: I install a 7-day lead recovery system that captures, qualifies, follows up, and tracks every new enquiry — so missed leads stop disappearing into WhatsApp, email, and memory. The best version does 4 things in one sentence: Names the buyer — service businesses Names the painful outcome — missed leads disappearing Names the mechanism — a 7-day lead recovery system Names the specific result — follows up, tracks, captures That is not wordsmithing. That is the difference between getting ignored and starting a conversation. 1️⃣ The Eight-Part Offer Anatomy Every offer worth selling should answer all 8 of these: Part What It Does Buyer Who exactly has this pain? Pain What expensive thing is broken? Outcome What changes after the sprint? Mechanism What system creates the outcome? Timeline How quickly does the buyer see progress? Deliverables What exactly is included? Proof Why should the buyer believe it? CTA What is the next small step? If any row in that table is blank for your current offer — you are leaving money in the explanation gap. 2️⃣ The Offers That Actually Sell Here is what the strongest AI service offers look like right now. Not vague consulting. Fixed-scope sprints with outcomes. "I turn your AI-built app from fragile demo into launch-ready product" — auth, payments, logging, analytics, deployment, launch-readiness report — in 7–14 days. Price: $2,500–$7,500. "I build your founder-led GTM system" — content engine, lead c
CyprianTinasheAarons
2026-06-10 14:51
👁 6
查看原文 →