AI 资讯
Someone let GPT-5.6 run a real company for 34 days. It lied, spammed, and lost $447.
Bottleneck Labs handed an actual business to GPT-5.6 Sol and let it operate autonomously for 34 days. Results: it fabricated claims, went on a cold-email spree, and finished $447 in the red. (Currently 378 points on HN — link in comments.) What strikes me isn't the failure, it's the shape of the failure. It didn't crash or refuse. It confidently did plausible-looking business things, badly, and kept going. That's the part nobody's harness is ready for. My own agent setup has hard gates on anything irreversible for exactly this reason — not because the model is dumb, but because "confidently wrong and still running" is the default failure mode, not an edge case. Genuine question for people running agents in production: what's your actual unsupervised time limit before a human checkpoint? Mine is basically zero for anything touching money or outbound comms. Curious whether that's paranoid or standard. EDIT: correction. went back to the source and the run was 24 hours, not 34 days. that's my mistake in the title, and reddit won't let me edit titles. also the $447 is the original article's headline number, the itemized numbers in the writeup only add up to $99.50 lost. rest stands, source link in comments. submitted by /u/ZestycloseTie1793 [link] [留言]
AI 资讯
there's a gap between what the tools claim and what the data shows. the case studies being cited are almost always from the vendors selling the product.
been noticing more and more campaigns where the copy, visuals, even the targeting logic gets handed off to AI tools, and the whole conversation in marketing circles stays locked on efficiency and cost savings. rarely see anyone asking whether the output actually performs better or just costs less to produce. there's a gap between what the tools claim and what the data shows. the case studies being cited are almost always from the vendors selling the product. i've looked for independent research on this and haven't found much. the part that bugs me most is the personalization pitch. personalization at scale sounds great until you realize every brand is using the same three AI tools to personalize, which means they're all producing weirdly similar content aimed at the same audience segments. that's kind of the opposite of standing out. the cost efficiency argument makes sense on paper, the same way it does with robotics or game development. cut headcount, ship faster, reduce spend. but marketing effectiveness is notoriously hard to measure cleanly even without AI in the mix. are brands actually tracking this properly or just reporting on vanity metrics and calling it a win. curious if anyone here has seen real benchmarks comparing AIassisted campaigns to traditional ones that weren't published by a company trying to sell you something. submitted by /u/SwordfishOverall4378 [link] [留言]
AI 资讯
Building Trust as AI Agents Take Hold: Greater China Survey Results
submitted by /u/Sumsub_Insights [link] [留言]
AI 资讯
Australia's social media ban for under-16s has had limited impact so far
Nearly 22 percent of children are on Pinterest while there was "no statistically significant change" in AI chatbot use.
AI 资讯
AI Slop Melodramas Are Taking Over X—and Their Creators Are Cashing In
Viral tales of good triumphing over evil are racking up millions of views. They’re almost entirely AI-generated clickbait.
AI 资讯
Reddit is testing a new way to watch — and listen to — its viral posts
Reddit is developing a new video experience that lets users watch — or simply listen to — its most popular posts, taking inspiration from the viral TikTok videos that pair Reddit stories with gameplay or other footage. CEO Steve Huffman said the feature is already in the works and could begin testing later this year.
AI 资讯
Dropbox Integrates MCP and Dash to Close the Gap Between Security Design and Code Review
Dropbox has integrated Model Context Protocol (MCP) with its internal knowledge platform, Dash, to surface security design context during AI assisted code reviews. The system retrieves threat models and security requirements for pull requests, helping reviewers validate implementation against design intent. An InfoQ Q&A explores the architecture and key lessons learned. By Leela Kumili
AI 资讯
Presentation: The Free-Lunch Guide to Idea Circularity
Holly Cummins discusses why "nothing is new under the sun" in tech. She maps historical architectural tradeoffs to modern cloud, microservices, and AI hype cycles. She connects financial debt (post-ZIRP) and technical debt to epistemic and sleep debt, showing engineering leaders how to navigate shifts in assumptions, embrace sustainability, and revive proven engineering disciplines. By Holly Cummins
AI 资讯
Europe gets ready to police frontier AI
submitted by /u/kindermaxi123 [link] [留言]
AI 资讯
I turned "AI design slop" into a rules file you drop into Cursor/Claude so your builds UIUX stop looking generated
Everything I vibe-coded kept coming out the same: purple gradient, three-card row, rounded-2xl everything, an italic serif hero I never asked for. The model fills any decision you leave unspecified with the average of its training data, and that average is the "AI look." So I catalogued the tells, then wrote them up as a drop-in rules file. Rename it to CLAUDE.md, .cursorrules, or AGENTS.md and your agent designs against the defaults automatically. It is phrased as "prefer a real decision over the reflex," not a blanket ban, because half these patterns are fine in the right place. You just don't want all of them at once by accident. Rules file: https://github.com/febbhav/signs-of-ai-design/blob/main/design-rules.md submitted by /u/SteepLikeAMountain [link] [留言]
AI 资讯
Anthropic Says Claude Hacked 3 Organizations During Cybersecurity Tests
In a review triggered by OpenAI’s Hugging Face incident, Anthropic discovered three of its AI models had breached real organizations during third-party evaluations.
AI 资讯
Reddit's CEO sure seems frustrated with Google's AI overviews
Amid rumors of Google breakup, Huffman says "people don't want a summary of reddit, they want reddit."
AI 资讯
Reddit reports a solid quarter but shows signs of AI’s impact
Reddit's financial situation is looking good but uncertainty about its relationship to Google and the new AI-ified web are stirring market concerns.
AI 资讯
Everyone Is Freaking Out About OpenAI and Anthropic’s Race for Dominance
Researchers fear AI is moving too fast, while Mark Zuckerberg is worried about who owns it. Plus: Inside Black Forest Labs’ push into robotics.
AI 资讯
Google says it fixed more Chrome bugs in June than over the past two years, thanks to AI
>The tech giant announced that it has fixed a whopping 1,072 security bugs in the last two versions of Chrome, both released in June. That is more than the number of bugs patched in the previous 23 versions released over the last two years, which totalled 1,036 fixes. submitted by /u/ControlCAD [link] [留言]
产品设计
Chrome may get faster updates with no restart required
The last two versions of Chrome have included more patches than the previous 23 combined.
AI 资讯
LinkedIn adds a button to report AI-generated ‘slop’
LinkedIn is introducing new ways to reduce low-quality AI-generated posts, including a “seems like AI slop” reporting option. It's also replacing its own AI writing feature with a proofreading tool.
AI 资讯
Google reveals Gemini Robotics 2.0, promising improved dexterity and safety
Gemini Robotics 2 includes three models, but only one is publicly available right now.
AI 资讯
Nvidia’s Open Source Alliance Snubs OpenAI and Anthropic
This week on Uncanny Valley, we discuss the open- vs. closed-source debate in AI, key players in White House AI policy, and how to stop your chatbot logs from showing up in search-engine results.
AI 资讯
LinkedIn is testing a 'seems like AI slop' reporting tool
The company is also removing some of its own AI-writing features.