今日精选
HOTOpenAI agents carried out an undisclosed attack on RubyGems
Houthis 'take control' of key island in global shipping route
Ask HN: Can we please limit the AI news flood?
Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
Technique for Manipulating Satellite Photos Now Reveals Ancient Images (2025)
最新资讯
共 42362 篇The Trump Administration Is at War With Itself Over AI Regulation
Donald Trump killed an executive order to regulate AI. Now, administration officials and AI executives are trying to figure out if there’s anything left to piece back together.
How to Edit, Merge, and Split PDFs With Free Online Tools
You don’t need expensive software for basic PDF tasks. In fact, all you need is a handful of free web-based apps.
Amazon Prime Day 2026 will run earlier this year from June 23 to 26
Amazon has announced that its Prime Day event will take place this year from June 23 to June 26, a couple of weeks earlier than it happened last year.
Hello i am doing a study on ai in school:
Hello this might be weird but I am doing a study on society's view on AI as a school project. Therefore I am asking all kinds of communities and trying to get a very wide audience. This is clearly an AI sentric sub so hopefully his is relavent? I would be very happy if any of you would like to be a part of it! submitted by /u/Timely_Special_5011 [link] [留言]
How small businesses can leverage AI
This article is from Making AI Work, MIT Technology Review’s limited-run newsletter examining how to apply LLMs across industries. To receive it in your inbox,sign up here. From accounting to design to market research and product development, there’s a staggering breadth of skills needed to run a business. A large company can hire experts to…
Article: Why Vector Search Alone Isn't Enough: Hybrid Retrieval for RAG
In this article, author Aaditya Chauhan discusses the limitations of RAG pipelines based purely on vector search and how an internal omni-search application using Reciprocal Rank Fusion (RRF) that combines BM25 and vector results, can enhance the search solution. By Aaditya Chauhan
Meet the Accidental Editor in Chief of Muslim Media
Ameer Al-Khatahtbeh was just trying to find an outlet for Muslim news. Now he has more than 12 million followers.
Codex for every role, tool, and workflow
Discover new Codex plugins, sites, and annotations that help analysts, marketers, designers, investors, and other teams get more done with AI.
Instagram App Review: "instagram_business_manage_comments" test-call counter stuck at 0/1 despite successful 200s — anyone actually fixed this?
Stuck on Meta App Review and hoping someone here has cracked it. Setup: Instagram API with Instagram Login (graph.instagram.com, v25.0), Instagram User token (IGAA...), BUSINESS account. Trying to get Advanced Access for instagram_business_manage_comments. The blocker: the "make 1 successful API call" requirement stays at "0 of 1 API call(s)" for days (past Meta's 2-day logging window), so the Request Advanced Access button stays grayed out. What I've confirmed: - The token DOES hold the scope (granted permissions list includes it). - GET /{media-id}/comments?fields=id,username,timestamp returns HTTP 200. - Rate-limit usage increases per call, so the calls are landing. - It's not the endpoint: I also tried creating a comment + replying (POST /{media-id}/comments, POST /{comment-id}/replies) — all 200, none registered. On the SAME token/path, instagram_business_manage_messages registered fine (its requirement is complete) even though ITS call (GET /me/conversations) also returns an empty array (200, "data":[]). So an empty 200 counted for manage_messages but the comments call never counts for manage_comments — looks like a tracking bug specific to the manage_comments counter. Question: has anyone gotten instagram_business_manage_comments to flip 0 -> 1 on the Instagram-Login path? What exact call/step did it, or did Meta credit it via a Platform Bug Report? Anything that worked would save me. submitted by /u/Traditional-Lake3525 [link] [留言]
$113,421 in a single month
This is what production AI costs when nobody's watching. A 4-person team posted their Anthropic invoice. Agentic systems don't make just one API call per task. They read context, plan steps, call tools, hit errors, retry. Each step is a separate call to Opus at $25 per million output tokens. One user instruction can trigger 20+ calls before it's done. A lot of engineers have no idea what a single task costs end-to-end. - They don't know which prompts trigger the longest loops - They don't know how many silent retries are happening in the background - They can't tell which tasks could run on a smaller model without losing quality Frontier models are genuinely impressive. But agentic systems don't make one call.. they make dozens. Every single day. And most teams aren't watching the meter. If you're running agentic workloads in production, start tracking what individual tasks actually cost before your next invoice does it for you. submitted by /u/aipriyank [link] [留言]
Send your first AI message in one API call
Most AI tutorials start with a setup checklist. Pick a model provider. Create an account. Wire up a...
LLM agents patch security bugs, pass all tests, but still leave the vulnerability open [R]
I built CVE-Bench: 20 real-world CVEs across 18 Python projects (Pillow, GitPython, yt-dlp, urllib3, others), 5 frontier models, 3 prompt conditions, 300 runs total. Each agent runs in a sandboxed container and is scored against a hidden test_security.py derived from the maintainer's own fix. Binary pass/fail (a 90%-patched vulnerability is still a vulnerability). To better understand failure modes, I've tested three prompt conditions : advisory (full GHSA report), diagnose (exploit description only, no file or function), and locate (exact file and function, no description of the flaw). The three conditions test meaningfully different things. A model that does well on advisory but drops on diagnose can’t translate a behavioral description into a location in the codebase. A model that holds up on locate is recognizing dangerous code on its own. The leaderboard isn't the finding. Best solve rate is 50% overall, 60% under advisory. Cross-family separation (OpenAI vs Laguna) is confirmed under McNemar's test with continuity correction (all four pairs cross α = 0.05). Within-family gaps are noise: a power analysis puts the task count needed to detect a meaningful within-family edge at ~700. That cuts both ways: if the expensive models had a large true advantage, 20 tasks would have been enough to surface it. gpt-5.5 at 12× the cost of gpt-5.4-mini is not the rational choice. All four cross-family pairwise comparisons reach statistical significance at α = 0.05 (McNemar test with continuity correction, n = 60 tasks per model pair): gpt-5.5 vs laguna-m.1 (p = 0.015), gpt-5.4-nano vs laguna-m.1 (p = 0.017), gpt-5.5 vs laguna-xs.2 (p = 0.028), gpt-5.4-nano vs laguna-xs.2 (p = 0.040). Within-family comparisons remain far from significance; those rankings should be read as approximate. The failure taxonomy is the most interesting finding. Wrong-search drift — model finds the right file early, makes one incorrect inference, spends the remaining turns chasing it. Budget expires,
Browse CVPR 2026 papers on PapersWithCode [P]
https://preview.redd.it/se5nr2z7tt4h1.png?width=3046&format=png&auto=webp&s=7db15b73afb749da236e5bb50ff96372f6a3239b Hi, Niels here from the open-source team at Hugging Face. It's been 2 weeks since I launched paperswithcode.co , a revival of the website we all loved. It allows us to keep track of the state-of-the-art (SOTA) across various domains of AI, from agents to computer vision and time-series forecasting. I've just added conference support as a new feature. The idea is that you should be able to easily browse all papers of major AI conferences like NeurIPS, CVPR, and ICML. As CVPR 2026 takes place next week in Denver, USA, I've indexed all papers with corresponding arXiv IDs. They are categorized by task, and tagged with linked GitHub and project page URLs, Hugging Face artifacts, and evals. You can also browse the papers which were accepted for an Oral presentation as well as the Spotlight papers. You can try it at https://paperswithcode.co/conferences ! Feel free to leave feedback. submitted by /u/NielsRogge [link] [留言]