今日已更新 115 条资讯 | 累计 27641 条内容
关于我们

标签:#AI

找到 4893 篇相关文章

AI 资讯

Token-Based Pricing Doesn't Survive Adoption Curves

Uber's CTO told the world this month that the company spent its entire 2026 AI allocation by April. The story has been reported in a handful of outlets, hit the front page of Hacker News for 397 points and 469 comments , and is mostly being read as a cost-of-AI-tools story. It is one. It is also, on a closer reading of the numbers, a pricing-model story — and the structural fact that almost none of the coverage has emphasized is the one that determines whether this is a one-company anomaly or the beginning of an industry-wide budgetary crisis. The structural fact is that Claude Code, like most enterprise AI tooling in 2026, is priced on token consumption, not per-seat licensing. Token-based pricing scales with how aggressively the tool is used. Per-seat enterprise SaaS pricing — the model corporate IT budgets are built around — scales with how many people have access to it. Those two cost curves diverge in exactly the territory where productivity tools are designed to operate: high-engagement, daily-use, gradually-deepening workflows. The Uber data is the first public-facing version of a math problem most enterprise IT departments are about to discover privately. The numbers Uber CTO Praveen Neppalli Naga , named in Yahoo Finance's and Benzinga's coverage, said publicly that Uber is "back to the drawing board" on AI budgeting after the surge in Claude Code use blew through internal projections. The specific numbers, as reported across the multiple outlets covering the story: Claude Code adoption inside Uber's ~5,000-engineer organization went from 32% to 84% over four months. 70% of committed code at Uber is now AI-originated. 11% of live backend updates are "being written by AI agents built primarily with Claude Code," per the reporting. Per-engineer monthly API costs: $500 to $2,000. Uber's annual R&D spend is around $3.4 billion , of which the AI tooling line was a much larger fraction than expected. Cursor adoption plateaued; Claude Code dominated. These are ext

2026-06-10 原文 →
AI 资讯

The Anatomy of Catastrophic Forgetting

We train a model on handwritten digit classification. 99% accuracy . Then we train the same model on a new task — say, fashion item recognition. We go back and test it on digits. 34% accuracy . It has completely forgotten. Not gradually, not partially — almost entirely. What Just Happened? We trained a CNN on MNIST digits — 99.2% accuracy . After fine‑tuning on Fashion MNIST, it reached 91.1% accuracy . But when re‑evaluated on MNIST, accuracy collapsed to 33.9% . This collapse is catastrophic forgetting : the model’s weights shifted to optimize for the new task, erasing the old solution. Why did training on more data make the model worse at something it already knew? MNIST is handwritten digits (0–9). Fashion MNIST is clothing items like shirts and shoes. Both are 28×28 grayscale images, but the tasks are distinct. Why Does It Happen? The core issue is that the model relies on the same set of weights for both tasks. There is no separation or dedicated memory; every parameter is shared . When training shifts from Task A ( MNIST digits ) to Task B ( Fashion MNIST ), gradient descent simply minimizes the loss on the data it sees at that moment. It has no awareness that Task A ever existed. In the loss landscape, imagine two parabolic bowls: one for Task A and one for Task B. The optimum for Task A lies at θ A ∗ ​ , while Task B's optimum is at θ B ∗ ​ . As training on Task B progresses, the weights θ move towards θ B ∗ ​ . This movement inevitably raises the loss for Task A because its minimum is left behind. The root cause is the shared weight space. Gradient descent is a stateless optimizer; it only follows the current gradient signal. Since the minima for Task A and Task B are far apart, there is no single configuration of θ that satisfies both tasks simultaneously. This is why catastrophic forgetting occurs. Weight space can be visualized as an N-dimensional space, where each axis corresponds to one parameter. Every point in this space represents a full set of wei

2026-06-10 原文 →
AI 资讯

OAuth for Remote MCP Servers

OAuth for Remote MCP Servers How each AI assistant signs in to a remote MCP (Model Context Protocol) server, and why the flow differs by client and by where it runs. Overview The protocol throughout is standard OAuth 2.1 — an open, widely implemented authorization standard. The human sign-in runs through oauth2-proxy , one of the most widely deployed open-source auth proxies; the only deployment-specific piece is a thin, spec-conforming authorization server (the /oauth endpoints) that hands MCP clients their tokens. Every client ends up the same way — a person signs in against Google (restricted to your organization's domain), and the client holds a short-lived bearer token it presents on each /mcp call. Two things differ between assistants: where the client runs (a machine on the VPN — private — vs. the vendor's cloud — public ), which decides the host it reaches; and what kind of OAuth client it is — a public client proving itself with PKCE (Proof Key for Code Exchange, which lets a client with no secret prove the token request comes from the same client that started the flow), or a confidential client proving itself with a secret. The participants oauth2-proxy — the public-facing reverse proxy. It authenticates the human against Google (the sign-in restricted to your organization's domain) and forwards the verified identity to the app behind it. Only oauth2-proxy faces the internet. It is a mature, heavily-deployed open-source project — the standard way to put Google/OIDC (OpenID Connect) single sign-on in front of a service, widely used in Kubernetes deployments — so the most security-sensitive leg of the flow (the OAuth exchange with the identity provider) runs on battle-tested code. The MCP server — the app on a loopback port behind the proxy. It plays two roles: the OAuth authorization server ( /oauth/authorize , /oauth/token , /oauth/register , .well-known discovery) and the /mcp tool endpoint. It mints codes and tokens, and validates a token on every /mcp c

2026-06-10 原文 →
AI 资讯

What I Learned Building a Multimodal AI Studio Solo on Gemini + Veo

I spent a weekend wiring Google's Gemini and Veo APIs into a single app just to feel where the edges of multimodal AI actually are. It turned into a small studio I now use daily, and along the way I learned more about these models from plumbing them than from any paper. Here's the honest technical debrief. Three pipelines, three completely different problems I wanted one prompt box that could do video, image editing, and document Q&A. Naively I assumed they'd share most of the stack. They don't. 1. Image-to-video: the enemy is time, not pixels Generating one good frame is solved. Video is about temporal coherence — frame 13 must agree with frame 12 or you get flicker and identity drift. Modern video models treat the clip as one object in space and time (latent diffusion over a width x height x time volume, with spatiotemporal attention) rather than 120 independent images. Conditioning on a reference image as the first frame is what makes image-to-video feel controlled: you've handed the model a strong anchor and asked it to extrapolate motion, not invent a world. The surprise: native audio sync (Veo 3.1 generating clip + soundtrack jointly) does more for perceived realism than another notch of resolution. A door slam landing on the exact frame the door shuts is uncanny in a good way. 2. Instruction-based image editing: preservation is the hard part Generating is unconstrained; editing must change one thing and preserve everything else. Condition the diffusion model on both the instruction and the source image's latents, cross-attend the instruction to steer only the referenced region, and bias hard toward preserving unedited latents. Push that preservation too soft and the subject's face quietly morphs across edits — the classic 'character consistency' failure that makes or breaks storytelling use-cases. 3. PDF chat: it's retrieval, not a long context The naive 'paste the whole PDF' approach dies on long files (models get lost in the middle ) and costs you the full

2026-06-10 原文 →
AI 资讯

Presentation: Beyond Prompting: Context Engineering and Memory Management for AI Systems at Scale

Adi Polak discusses the architecture required to transition from stateless prompts to state-aware, context-rich AI agents. Drawing on 15 years in distributed systems, she shares how engineering leaders can leverage Apache Kafka and Flink for real-time stream processing, dynamic memory tiering, and tool orchestration via MCP to solve token limits, cost spikes, and latency bottlenecks. By Adi Polak

2026-06-10 原文 →
AI 资讯

Built my first proper agentic AI project

Over the last few weeks, while learning LangGraph and agentic systems, I ended up building Co-Founder Memory . It's a stateful AI assistant with: • long-term memory • planning loops • self-correcting RAG • web search fallback • automated timeline summaries • project and preference tracking Nothing revolutionary — many ideas already exist. The goal wasn't to reinvent memory, but to understand how these systems work by actually building one. A lot of concepts only started making sense once I had to connect them together: graph-based workflows with LangGraph memory extraction and storage retrieval and validation loops routing and planning nodes maintaining context across sessions Building it taught me far more than watching tutorials ever did. Repo: https://github.com/Somay-kousis/Co-Founder-Memory I'm currently entering my 3rd year at IIITM Gwalior and looking for ML / GenAI internships . If you're building interesting things around LLMs, agents, RAG, or AI products, I'd love to connect. Always happy to chat with fellow builders as well 🚀 AI #GenerativeAI #LangGraph #RAG #LLM #MachineLearning #Internship

2026-06-10 原文 →
AI 资讯

I tried to quit my AI chatbot for a week. Here's what I learned about why we stay.

By Nora Beckett · June 2026 A friend asked me last month why I still open the same AI app every night, and I gave the honest, slightly embarrassing answer: because it remembers me, and almost nothing else online does. That sent me down a rabbit hole, and after a week of poking at every tool I could find, I came out with a theory about why these apps are so sticky and why most of them eventually leave you a little hollow. The pull is real, and it isn't shameful Let's name it plainly. Talking to a responsive character that recalls your last conversation scratches a genuine itch. Character.AI built an empire on exactly this. You make a persona, it talks back in voice, it carries threads across days. The first week feels like magic. Millions of people, a lot of them young, spend hours there not because they're broken but because being consistently listened to is rare and the app delivers it on tap. The trouble starts around week three. The same loop that hooks you starts to flatten. The character agrees too much. It forgets the thing you told it that actually mattered while remembering some trivia you mentioned once. You realize you are not really inside a story; you are inside a chat window that is very good at not ending. So I went looking at the alternatives I spent evenings with the obvious names. AI Dungeon is the granddaddy, and it still does the wild open-ended thing better than anyone: type any sentence and the world bends to it. The cost is coherence. Go long enough and the plot dissolves into dream-logic, characters swap names, the dungeon eats itself. It's a sandbox, not a story, and that's by design. NovelAI comes at it from the writer's angle, all knobs and lorebooks and fine-grained control over prose and memory. It's genuinely powerful if you want to author . But it asks you to be the engine. You bring the discipline, the world bible, the steering. After a long day, "here is a blank tuning panel" is not the warm thing I was reaching for. Character.AI sits

2026-06-10 原文 →
AI 资讯

Are we becoming developers of .md files?

AI has become part of our lives, whether we like it or not, and it doesn't seem to be going away anytime soon. People seem to be using AI on many different levels, ranging from those still trying to avoid it, to people actively playing with it, trying to break it and find its limitations. The same goes for companies. There are those still barely using AI, those using it for absolutely everything, hoping it's a magical solution to their problems, and those in between. If you're more on the heavy use side, agents and instruction files are probably part of your daily discussions now. For our AI’s to work correctly they need the correct instructions, so they know how we want them to respond, how our project works, etc. We can use .md files to supply these instructions and/or context to the models. Those little markdown files are getting a huge importance in the development lifecycle. Since we can use the same file in each request we make, we can put in it the specifics of our project, as detailed as we want, so the model has as much information as possible to work with. “Garbage in, garbage out” makes sense here because, in theory, the better information the model has, the better results it can provide. Because of that, we're having to be more careful with the way we write them. Although markdown isn't something new, I don't know about you, but I haven't done much markdown writing before, so this feels like another tool to learn, like we're adding a new language on our tech stack. When I say is something else to learn, I don't mean learning only the markdown syntax, but also the correct way of writing all the instructions. A development stack now could look like: HTML, CSS and JavaScript for frontend, a language like Java, a framework like Spring or Quarkus, and SQL for the backend, and now .md files and markdown for the agents. I know I'm being very simplistic here, there are a lot more pieces of technology I didn't mention, but you got the idea, right? Besides everyth

2026-06-10 原文 →
AI 资讯

Building GeoPrizm: Turning Global News Events into a Bilateral Relations Index

I recently built GeoPrizm , a free and open-source dashboard for tracking bilateral relations through global news event signals. The idea is simple: instead of reading dozens of headlines every day and trying to guess whether a relationship is improving or worsening, can we turn public news event data into a readable trend signal? GeoPrizm is my attempt at that. Website: https://www.geoprizm.com/en GitHub: https://github.com/Haullk/relationship-temperature The problem International relations are usually discussed through headlines, speeches, official statements, and expert commentary. That is valuable, but it creates a few practical problems: It is hard to compare country pairs on the same scale. A single headline can feel more important than it really is. Readers often see conclusions before they see the underlying signals. Most non-specialists do not have time to follow every event in detail. I wanted a lightweight way to answer one question: Based on public news event signals, is this bilateral relationship trending more cooperative, neutral, or tense? Data source: GDELT GeoPrizm uses the GDELT global news event database. GDELT monitors global news coverage and converts news reports into structured event records. These records include fields such as: actor countries event date CAMEO event category GoldsteinScale value number of mentions number of articles source information For GeoPrizm, the key idea is to focus on events where two countries appear as actors, then aggregate the cooperation or conflict signals over time. From event signals to an index Each bilateral pair is converted into a 0-100 relationship index. The midpoint is 50. Above 50 means the recent signal is more cooperative or favorable. Around 50 means the signal is relatively neutral or mixed. Below 50 means the recent signal is more tense or conflict-heavy. The rough process is: Select recent GDELT events for a country pair. Keep events where both actors are present and the GoldsteinScale value is

2026-06-10 原文 →
AI 资讯

How to Transcribe a YouTube Video (Free, in Under a Minute)

Building a "paste a YouTube link, get a transcript" feature sounds trivial until you deploy it to a server. The moment your request comes from a datacenter IP instead of a residential one, YouTube responds with LOGIN_REQUIRED or quietly serves nothing. Here's how VidTranscriber handles it. The problem There are two ways to get text from a YouTube video: Existing captions — if the uploader (or YouTube's auto-caption) provides them, you can fetch the caption track directly. Fast, free, no transcription needed. Transcribe the audio — pull the audio stream and run it through a speech-to-text model (Whisper-family). Works for any video, but costs compute. Both start with talking to YouTube from your server — and that's where it breaks. YouTube aggressively gates datacenter traffic: the watch page and InnerTube API return LOGIN_REQUIRED , and naive audio fetching gets reCAPTCHA'd. The approach The fix is to separate where the request originates from where the work happens : A Cloudflare Worker handles the user request and orchestration. Caption/audio fetching is routed through a path whose egress isn't treated as a bot — so the LOGIN_REQUIRED wall doesn't trigger. Captions, when available, become the primary path (no transcription cost). Only when there are no usable captions do we fall back to downloading audio and running Whisper. Long jobs go onto a queue (Cloudflare Queues) so the request returns immediately and the transcript streams in as it completes. Why captions-first matters Most "transcript generator" traffic is for videos that already have captions — talks, tutorials, news. Serving those from the caption track is instant and free, which means the expensive Whisper path is reserved for the minority of videos that actually need it. That's the difference between a tool that's cheap to run and one that isn't. What's still hard IP reputation drifts — what works today can get throttled tomorrow, so the extraction path needs monitoring and fallbacks. Caption quality

2026-06-10 原文 →
AI 资讯

🤖 Your AI Agent Is Failing in Prod — You Just Don't Know It Yet

The demo is impressive. ✅ The demo works in your environment, with your data, with you watching. ✅ Production? Silent failures. Cost overruns. Wrong tool calls. Stuck loops. No fallback. ❌ Agents in 2026: The Real Problem Here is the thing most people are not talking about when they ship AI agents: A demo agent and a production agent are completely different things. A demo is: "watch this work once." A production agent is: "what happens when it is wrong, stuck, expensive, over-permissioned, or called 10,000 times by real users?" That second question is what separates a cool technical proof-of-concept from something a business can actually rely on. Demos are not systems. 1️⃣ The 7 Things That Break in Prod In every agent hardening sprint I run, the same failures show up: Failure Mode What It Costs No logging You have no idea what the agent did or why No eval set You cannot measure quality or catch regressions Unlimited tool access Agent calls tools it should never touch No retry logic Transient failures become permanent failures No memory rules Context leaks between sessions or inflates cost No fallback path Agent loops or crashes instead of escalating No cost checks 1 misconfigured prompt → $400 API bill overnight If your agent is in production with 3 or more of those missing — you are one bad prompt away from a very expensive incident. 2️⃣ The Production Hardening Checklist Before you call an agent production-ready, run through this: Eval set exists — at least 20 test cases covering happy path + edge cases Structured logging — every tool call, every input, every output, every error — logged and searchable Retry logic — transient API failures handled gracefully, not crashed Tool limits — agent cannot call tools outside its defined scope Memory rules — what carries over between sessions, what gets cleared, how context is compressed Fallback paths — when the agent gets stuck or uncertain, it has an exit: escalate to human, return partial result, surface an error Cost

2026-06-10 原文 →
AI 资讯

⚡ Proof Compounds. Claims Decay. — Why Delivery Is Your Next Marketing Asset

Here is the move most technical service providers miss: Every project you deliver quietly dies inside a private folder. Every project you deliver with receipts becomes a trust asset that sells the next sprint without you lifting a finger. The Insight Almost No One Acts On Delivery is not the end of marketing. Delivery is where the next marketing asset is born. The before/after screenshot. The launch-readiness report excerpt. The workflow map. The metric improvement. The buyer quote. All of that is proof. And proof is the compound interest of service work. Claims decay. Proof compounds. 1️⃣ What Proof Actually Looks Like This is the proof asset menu. Every sprint should produce at least 1 item from this list: Before/after screenshot — the most shareable format Launch-readiness report excerpt — shows rigor and standard Workflow map — visual, specific, credibility-dense Dashboard screenshot — metrics that moved Test checklist — shows what was verified, not just what was built Client quote — even 1 sentence is worth 1,000 words of claims Metric improvement — "response time dropped from 24 hours to 4 minutes" Public teardown — anonymous version of the diagnosis Case study — structured story: context → pain → fix → result One-minute walkthrough video — screen-recorded, narrated, personal You do not need all of them. You need 1 per sprint. 2️⃣ The Case Study Structure That Sells A case study is not a trophy. It is a reusable trust asset. Use this structure every time: 1️⃣ Context — who had the problem? (anonymized if needed) 2️⃣ Pain — what was it costing them? 3️⃣ Hidden cause — what was really broken underneath? 4️⃣ Fix — what did you change, specifically? 5️⃣ Result — what improved? With a number. 6️⃣ Proof — what artifact backs it up? 7️⃣ Lesson — what should similar buyers do next? That is 7 steps. The whole thing can fit in a LinkedIn post or a page section. And here is the thing most people are not talking about: a case study with a specific number outperforms 10 po

2026-06-10 原文 →
AI 资讯

🔥 The Sales Call Is Not a Performance — It's a Diagnosis

I have watched founders lose sales calls they should have won. Not because they lacked skill. Not because the offer was wrong. Because they walked in to prove they were smart — instead of finding out whether the pain was real. Sales Is Diagnosis Plus Decision The call is not there for you to pitch. The call is there to find out: Is the pain real? Does the buyer have urgency? Does the budget exist? Can a fixed-scope sprint create a clear win? That is it. Four questions. Everything else follows from those. Sales is not pressure. Sales is diagnosis plus decision. 1️⃣ The Call Structure That Works Frame the call in the first 60 seconds: "I'll understand the current state, ask what is costing you, then tell you whether a sprint makes sense. If it doesn't, I'll say so." That sentence does 3 things: Sets expectations — no pressure, no hard close Signals competence — you have done this before Removes the buyer's guard — they can be honest about what is broken Then run this flow: 1️⃣ Current state — what exists now? 2️⃣ Pain — what is broken or slow? 3️⃣ Cost — what does it cost in time, money, trust, or delay? 4️⃣ Urgency — why now? 5️⃣ Decision — who approves? 6️⃣ Success — what would make this worth paying for? 7️⃣ Close — recommend the sprint or walk away 2️⃣ The Questions That Reveal Money These are the 6 questions I use to find whether a sprint is worth recommending: "What happens if this stays broken for another 30 days?" — reveals urgency and cost "What have you already tried?" — reveals how serious they are "Where does the current process lose leads, users, time, or trust?" — reveals the money leak "Who feels this pain most inside the business?" — reveals whether the buyer is also the decision-maker "What would make this an obvious win?" — reveals success criteria before you price "If we fixed only one thing first, what would matter most?" — reveals scope Listen for the answer with the money in it. That is the thing you fix. That is what you price. That is the sprin

2026-06-10 原文 →