今日已更新 381 条资讯 | 累计 42058 条内容
关于我们

标签:#ai

找到 7617 篇相关文章

AI 资讯

Marca d'água em textos gerados por IA

Introdução A Anthropic anunciou recentemente a inclusão de uma marca d'água nos textos gerados pelos modelos Claude. O objetivo é distinguir conteúdos gerados por humanos daqueles criados por IA generativas. Funcionamento O prompt enviado é convertido em tokens que são usados para calcular a probabilidade do próximo token , se repetindo até uma resposta que faça sentido seja retornada para o usuário. Esse é o processo padrão utilizado pela maioria dos modelos de IA Generativa. Agora a Anthropic adotou uma abordagem determinística para selecionar o próximo token a partir da lista de prováveis candidatos, que continua sendo gerada aleatoriamente. Esta nova abordagem usa uma chave privada e parte do contexto já gerado para selecionar o próximo token, sucessivamente até a geração total do texto que o usuário recebe como resposta ao prompt inserido. De acordo com a Anthropic, a adição desse identificador não consumirá tokens adicionais nem tornará o modelos mais lento para responder. De forma resumida, o diagrama a seguir mostra o funcionamento dessa abordagem. O que motivou esta ação Em 2024 a União Europeia aprovou Lei de Inteligência Artificial , primeira lei criada para regulamentar a inteligência artificial, prevista para entrar em vigor a partir de 02/08/2026. O Artigo 50 n.º 2 determina que: "..., incluindo sistemas de IA de finalidade geral, que geram conteúdos sintéticos de áudio, imagem, vídeo ou texto, devem assegurar que os resultados do sistema de IA sejam marcados num formato legível por máquina e detectáveis como tendo sido artificialmente gerados ou manipulados." . 1 Algumas das Big Techs assinaram o pacto de adesão e estão implementando a parte técnica de acordo com cronogramas próprios. Entretanto uma rápida pesquisa mostra que a aplicação prática por parte das empresas proprietárias de modelos LLM ainda é pequena. Entre os motivos citados estão: pode tornar os textos rastreáveis; perda estimada de até 30% dos clientes da plataforma fragilidade técnica

2026-09-07 原文 →
AI 资讯

Your LLM Trace Is Green. Why Is the RAG Answer Still Wrong?

TL;DR Many LLM observability setups capture prompts, outputs, tokens, and latency while leaving retrieval failures hidden. A single search call may conceal query rewriting, filtering, fetching, deduplication, reranking, and evidence selection. A useful trace connects the original question to the effective query, returned sources, selected passages, and final claims. Retrieval tracing helps distinguish missing, stale, or ignored evidence from a genuine generation failure. Production teams should measure freshness, duplicate evidence, citation coverage, and cost per grounded answer. A user asks your AI assistant whether a product still supports a particular feature. The assistant responds confidently and links to the company’s documentation. The model request succeeded. Latency was normal. Token usage stayed within budget. No tool call failed. Every indicator on the dashboard is green. The answer is also six months out of date. The model trace cannot tell you whether the system searched for the wrong phrase, preferred an old page, discarded a better result, or ignored the correct evidence. It only shows the context that eventually reached the model. That is the blind spot in model-centred observability. For RAG applications and web-connected agents, the useful unit of observation is not the model call. It is the complete evidence path. A Successful Model Call Can Still Be a Failed Request A typical LLM trace records the prompt, response, model name, token consumption, latency, errors, and perhaps a tool invocation. That is useful for diagnosing slow requests, malformed inputs, and unexpectedly expensive generations. It does not tell you whether the model received the right facts. In a retrieval application, the final prompt is assembled by an upstream system. That system may rewrite the query, choose a search provider, apply time or domain filters, fetch pages, extract text, remove duplicates, rerank candidates, and select passages for the context window. The model ca

2026-09-07 原文 →
AI 资讯

Pantrybridge

This is a submission for Weekend Challenge: Generosity Edition I wanted to build something for this challenge that didn't just talk about generosity but actually meant something, and felt beneficial. This tool can make it easier to go from "I have food to donate" to "I'm donating food". You take a picture of your pantry shelf, Gemini figures out what's actually in it, and the app turns that into a recipe for whoever receives it, a handwritten-style note of kindness, a real way to find a food bank near you, and a printable manifest to hand over at drop-off. What I Built PantryBridge is a small AI-powered toolkit for food donation. The flow is: 1. Scan your pantry. Upload a photo (or pick one of three one-click sample hauls if you don't have a pantry photo handy). Gemini does multimodal image analysis and returns a structured inventory: item names, categories, estimated quantities, dietary tags, urgency, and packaging condition. ( Sorry, GIPHY messed up my gif ) 2. Review the inventory. Everything shows up in a clean table with donation-readiness stats and a volunteer tip generated specifically for that haul. (If you need to, you can delete or add items!) 3. Find a real drop-off location. Enter your zip code and the app confirms your city/state (via a real geocoding lookup) and links you straight to Feeding America's actual food bank locator, so you're finding a real place to donate, not a mock one. 4. Get a recipe and a kindness note. Gemini writes a short recipe using mostly what you're donating, plus a genuinely warm, non-patronizing note to include with the box. 5. Print a donation manifest. A little printable card with the itemized contents and a mock barcode/QR for quick intake logging, with confetti when you pledge or print. Demo Try the Live App If you'd rather run it yourself: git clone https://github.com/780s/pantrybridge.git cd pantrybridge npm install npm run dev Drop a GEMINI_API_KEY into .env.local to hit the real Gemini API. Without one, every route qui

2026-09-07 原文 →
AI 资讯

HANDOFF: Give the Appliance. Pass on the Know-How.

This is a submission for Weekend Challenge: Generosity Edition A donated washer can reach its next home with everything it needs — except the one thing a manufacturer manual cannot contain: what happened to this specific machine. The person who repaired it knows what was replaced, what was tested, how this unit should be started, and what was packed with it. The recipient usually does not. That gap is what HANDOFF carries. Give the appliance. Pass on the know-how. HANDOFF lets a refurbisher speak once, then turns that short, item-specific explanation into a bilingual voice-and-text handoff that stays with the appliance through one durable QR tag. The volunteer already has the knowledge in their head. Speaking for 20 seconds is cheaper and more natural than writing custom instructions, translating them, formatting them, and printing them. And the recipient should not need an account, an app, or an English-first interface just to understand the thing they were given. What I Built HANDOFF is an object-specific knowledge handoff for donated and refurbished equipment . A refurbisher records a short voice note about the actual appliance in front of them. HANDOFF then: cleans the real recording with ElevenLabs Voice Isolation creates an English ↔ Spanish voice handoff with ElevenLabs Dubbing v2 retrieves readable source and translated text persists the completed media gives the handoff one durable ID generates a printable QR tag that travels with the appliance What the recipient gets The recipient sees their language first. For the verified English → Spanish sample: Español — Recipient English — Original They can play the recipient-language voice, read the same handoff as text, and switch both audio and text back to the original together. If the audio cannot load, the readable handoff remains available. Scan. Listen or read. The technician workflow is deliberately small: record → clean + dub → attach No recipient profile. No manual translation step. No long form. Why this

2026-09-07 原文 →
AI 资讯

Seattle Times and Newsday sue OpenAI and Microsoft for infringement

The Seattle Times and Newsday are just the latest plaintiffs to take OpenAI to court, alleging copyright infringement. The two outlets say the company used their journalism as training data for its AI models without permission and often reproduces passages from their reporting in response to user queries. This is similar to lawsuits filed by […]

2026-09-07 原文 →
AI 资讯

You Aren't Choosing an AI Tool. You're Choosing Who Gets Paged at 2 AM.

Part of AI Leadership in the Real World — how leaders turn scattered pilots into governed, adopted, measurable capability. TLDR: A support agent doing 50,000 chats a month needs ~3.5 FTE and $500k+/year just to stay accurate — while a typical 100-seat Copilot rollout sees only 20-30 seats used weekly. For SMBs, build-vs-buy isn't about features. It's about what you can afford to own for 24 months. We thought we were choosing a tool. We were really choosing a future dependency, a support queue, a governance burden, and a second bill that arrives a year later. Every vendor demo promised acceleration, control, and simplicity at once. Every internal proposal promised flexibility, ownership, and leverage. Nobody said both bills arrive late — one in engineering on-call, the other in consumption meters. Good platform decisions feel a little boring at first and very smart a year later. Why AI is special (and why old build-vs-buy math breaks) Traditional software mostly stays still when you leave it alone. AI doesn't: It drifts. Knowledge changes, customer language shifts, users ask harder questions once they trust it. Accuracy quietly drops from 90% to 70% with no error log. It speaks for you — legally. A wrong Confluence page is embarrassing. A wrong chatbot answer is a commitment a tribunal can enforce. It lives on someone else's deprecation clock. OpenAI gives at least 6 months before retiring a GA model. That's a hard deadline, not a backlog item. Prompts, evals, and output parsers all need rework. It multiplies cost per request. One human click = one action. One agent resolution = 6 lookups, drafts, updates, and logs — each potentially metered. It turns connectors into permanent work. Salesforce, SharePoint, Jira, Zendesk all change auth, rate limits, and APIs. Your agent keeps running while its knowledge goes stale. Gartner predicts 40%+ of agentic AI projects will be canceled by end of 2027 on cost, unclear value, and weak risk controls. McKinsey's State of AI 2025 (

2026-09-07 原文 →
AI 资讯

Handover: small charities know what hurts, not what skill they are missing

This is a submission for Weekend Challenge: Generosity Edition What I Built Handover takes a plain description of what is going wrong inside a small charity and works out the role that would fix it. Not the role they asked for. The one they actually need. You type something like "our books are a mess, and we have missed two filing deadlines". It comes back with a full trustee role: the diagnosis, what the person would do, a deliberately short list of essential skills, an honest time commitment, and an advert you can paste straight into your newsletter. Then a volunteer pastes their CV, badly, and gets scored against every open role with a reason and an honest note on where the fit is thin. Why A couple of days ago I got an email saying Reach Volunteering is closing after 45 years. It genuinely hurt to read. Reach connected small UK charities with people who wanted to give them professional skills. Last year it placed 5,996 volunteers and trustees across 2,440 organisations. The people it placed contributed around £60 million in expertise. Ninety-six per cent of those organisations ran on under £1 million a year, and nearly half on under £50,000. It is not closing because the work stopped mattering. It is closing because funding for the infrastructure that helps small charities build capacity has dried up. Reach was the largest single source of trustees in the sector, and it is shutting at the peak of its impact. I volunteer as a digital navigator, which mostly means sitting with people who have been handed a system that assumes a confidence nobody ever gave them. You watch someone decide they are the problem, when the thing in front of them was just badly built. Reach existed to stop small charities from feeling like that about their own gaps, and now it is shutting down. I cannot rebuild 45 years of relationships in a weekend. So I picked the one piece of what Reach did that was pure expertise rather than headcount, and rebuilt that. The thing everyone gets wrong E

2026-09-07 原文 →
AI 资讯

Local Embeddings vs. API Embeddings — Why I Chose sentence-transformers

Every RAG pipeline needs to convert text into vectors. The question is where that conversion happens. You have two options: run an embedding model locally on your own hardware, or call an API that runs the model on someone else's hardware. Both work. The right choice depends on your constraints — and understanding the tradeoffs is more useful than a recommendation. This article is about why I chose local embeddings with sentence-transformers/all-MiniLM-L6-v2 for this pipeline, and when I'd switch to an API. What Embeddings Actually Do Before the tradeoffs, a quick grounding on what's happening. An embedding model takes text and converts it into a fixed-size vector of floating-point numbers — a list of 384 numbers in the case of all-MiniLM-L6-v2 . That vector encodes the semantic meaning of the text in a way that allows mathematical comparison. Two pieces of text with similar meaning produce vectors that are close together in the 384-dimensional vector space. "Authentication failed" and "login was rejected" are semantically similar — their vectors will be close. "Authentication failed" and "quarterly revenue report" are semantically distant — their vectors will be far apart. This is what makes retrieval work. When you embed a query and search for the nearest chunks, you're finding chunks that are semantically similar to the question — not just chunks that contain the same keywords. The embedding model determines the quality of this semantic matching. A better model produces vectors where semantic similarity maps more accurately to vector proximity. The Local Embedding Choice My pipeline uses sentence-transformers/all-MiniLM-L6-v2 via ChromaDB's SentenceTransformerEmbeddingFunction : from chromadb.utils.embedding_functions import SentenceTransformerEmbeddingFunction embedding_fn = SentenceTransformerEmbeddingFunction ( model_name = " sentence-transformers/all-MiniLM-L6-v2 " ) This runs entirely on your local CPU. No API key, no network request, no cost per embedding,

2026-09-07 原文 →
AI 资讯

Your prompt system has no tests, and that is why you cannot tell it is broken

Tags: ai , python , testing , showdev Code fails loudly. A prompt system fails in silence, and it fails while still producing something that looks completely fine. I found this out the slow way. I had built a multi-skill agent system: 15 skills, nine commands, each one writing structured JSON that the next one reads. It worked for weeks. Then it did not, and I could not tell you when it stopped, because nothing ever threw. A skill quietly stopped writing one field. The next skill read a null and carried on. The final score came out a few points off, in a document that read exactly as convincing as it had the week before. Plausible output is the one thing these models are never bad at. That is precisely the problem. What a test even means here You cannot assert on the prose. Run the same prompt twice and you get different words, and that is fine, because the words are not the contract. Something else is. Three things turned out to be testable, and together they catch nearly everything: The arithmetic. My system scores six weighted dimensions and applies a penalty when any dimension falls below a floor. That is deterministic. The model produces the dimension values, but the final number is a function of them, and a function is something a checker can recompute without going anywhere near the model. If the number on disk disagrees with the number the checker computes, one of them is lying and it does not matter which. The shape. Every skill writes a file with an expected structure. Required fields, enums for anything constrained, explicit nullability, conditional requirements where one field's presence forces another. This is schema validation, and it is unglamorous, and it caught more real regressions than anything else I wrote. The prose rules that are actually numbers. A memo has to sit inside a word budget. It has to contain its required sections. It has to cite at least three URLs that are shaped like URLs. None of that judges quality, and all of it catches drift,

2026-09-07 原文 →
AI 资讯

The Overhead Ratio Is Lying to You — I Built an AI Tool to Prove It

This is a submission for Weekend Challenge: Generosity Edition What I Built GlassPocket — a tool that argues against the "overhead ratio," the dominant heuristic people use to judge charities (what % of donations go to "programs" vs. "overhead" like staff and infrastructure). That heuristic punishes exactly the investment that makes a charity effective, and it drives what nonprofit finance people call the "starvation cycle" — orgs under pressure to look lean end up under-staffed and under-resourced. You search a US 501(c)(3), and instead of a single overhead percentage, GlassPocket pulls their IRS Form 990 history (via ProPublica's Nonprofit Explorer API) and shows: Reserve months — how long the org could run on savings alone (low reserves = fragile, not "lean") Operating margin trends across up to 13 years of filings Staff-investment share — reframed as capacity, not waste Fundraising cost per dollar raised — a narrower, more honest efficiency metric than the classic ratio A peer-percentile chart against ~70 similar organizations in the same category A Gemini-written "myth-buster" card pairing each overhead-ratio assumption with what the numbers actually show A grounded chat box — ask follow-up questions about that specific org's finances, answered only from its own filing data Demo Live app: https://glasspocket.vercel.app/ Code hassan-2050 / glasspocket Overhead-ratio myth buster for US charities — Form 990 data via ProPublica, Gemini narrative generator GlassPocket — Overhead Myth Buster Live: glasspocket.vercel.app Category: Overall Winner + Best Use of Google AI (Gemini-powered narrative generator and chat) The Hook Most charity-rating tools reinforce the harmful "overhead ratio" myth. This contrarian tool argues against that dominant heuristic by reframing efficiency around outcomes and reserves. What It Does You search a US charity by name, and it pulls their IRS Form 990 history to generate a plain-English context-aware financial explainer that debunks the o

2026-09-07 原文 →
AI 资讯

How Freebuff, AgentRouter, OpenRouter, and Experiential Labs Give You Free AI Models (And the Business Tactics Behind It)

Frontier AI models are expensive to call directly. A single day of heavy Claude or GPT-5 usage in an agentic coding loop can rack up real money. But a small cluster of gateways and coding-agent products has figured out how to hand developers meaningful free access anyway. This post breaks down four of them — Freebuff, AgentRouter, OpenRouter, and Experiential Labs — and the actual tactics each one uses to keep the lights on while giving inference away. 1. OpenRouter — the "free router" and community-subsidized models OpenRouter is a unified, OpenAI-compatible API that sits in front of hundreds of models from dozens of providers. Its free tier isn't a special OpenRouter model — it's a curated set of models, mostly open-weight ones like DeepSeek R1, Llama variants, and Qwen releases, that carry a literal $0/M-token price tag because providers or OpenRouter itself are subsidizing the compute. The tactic: instead of making you pick a free model by hand, OpenRouter built openrouter/free , a router that automatically picks a working free model for each request, smart enough to filter for whatever the request needs — image understanding, tool calling, structured outputs, and so on. That's a neat trick: it turns "which free model works today" from a research chore into a solved problem, since free-model availability shifts constantly and the router absorbs that churn for you. To keep this sustainable, OpenRouter caps usage per key — community trackers put it at roughly 20 requests per minute and 200 requests per day on the free tier — and openly frames free access as ecosystem-building: it says free models help democratize access to AI and let large numbers of people experiment and learn, while it keeps expanding capacity by onboarding new providers and covering some costs directly. In plain terms, the free tier is marketing and community goodwill; paid usage across the rest of the catalog is the actual business. Using it is as simple as pointing any OpenAI-compatible SDK a

2026-09-07 原文 →
AI 资讯

The Hook System — Blocking AI Mistakes with Structure

This is chapter 4 of my book **Building Autonomous AI Agents with Claude Code * — a field guide to turning Claude Code from a coding assistant into an agent that remembers, verifies its own work, and knows when to stop. Everything below is from a system I actually run every day on one Windows PC.* 1. A Hook Is a Safety Mechanism Outside the AI A rules file is something the AI tries to follow ; a hook is something the system uses to make it be followed . This difference is bigger than it looks. Rules get buried as context grows longer, get skipped when things are urgent, and "just this once" exceptions pile up. Hooks don't do that. Point Timing Typical use UserPromptSubmit Right after the user types input Automatic context injection (record summaries, related rules) PreToolUse Right before a tool runs Blocking dangerous actions (gates) PostToolUse Right after a tool runs After-the-fact checks (contamination detection, follow-up procedure reminders) Stop When the response ends Quality gates (forbidden-word detection, verification requirements) Registration happens in one place, the settings file. { "hooks" : { "PreToolUse" : [ { "matcher" : "Write|Edit" , "hooks" : [{ "type" : "command" , "command" : "python C:/hooks/record_gate.py" }] } ] } } 2. Pattern A — The Blocking Hook (Gate) This is a gate that blocks "attempts to modify a file without reading the records first." What follows is a shortened version of one actually in use. import json , sys , time from pathlib import Path STATE = Path ( tempfile . gettempdir ()) / " read_state.json " REQUIRED = [ " memory/diary.md " , " memory/mistakes.md " ] payload = json . load ( sys . stdin ) # hooks receive the tool call on stdin tool = payload . get ( " tool_name " , "" ) if tool == " Read " : state = json . loads ( STATE . read_text ()) if STATE . exists () else {} state [ payload [ " tool_input " ][ " file_path " ]] = time . time () STATE . write_text ( json . dumps ( state )) sys . exit ( 0 ) state = json . loads ( STA

2026-09-07 原文 →
AI 资讯

I let software run my station for 8 weeks, unattended. What I learned.

Hi, my name is Yaniv Morozovsky, and I have worked in the radio industry for more than three decades. Most of that on SAM Broadcaster and, lately, AzuraCast. On 13 July I did something I had wanted to try for a long time: I handed one of my internet stations to automation completely, with nobody in the studio, and did not touch it. It is still on air today: https://ystream.live Some honest notes for anyone who runs a station and has wondered about the same. Segues are the whole game. The thing that makes automation sound like a jukebox is the fade timer. Every song fading at the same point, every cold ending faded when it should stop dead. What fixed it was analysing every song once (intro end, vocal entry, outro, the real ending) and crossing on those points per song. Once that was in, listeners stopped noticing there was no one there. Making an automated station sound like a DJ is mostly making the crossfades right, and that was our biggest achievement from day one. The voice matters less than where it lands. I cloned my own voice for the breaks, and honestly the voice itself is not what people comment on. What they hear is whether the talk-up ends before the vocal. When it lands, it sounds like radio. I let the AI DJ open every hour, read the weather, and in the last week also read a short news segment when the hour begins. Scheduling by rules beats scheduling by clocks, most of the time. Key, tempo, energy, era, genre, artist separation, and a few written rules like "start every hour with an international song". The hour writes itself and I stopped building clocks. What I gave up was the twenty years of Clockwheel habits, and there were days I missed them. Some of the most annoying bugs of automation software, like repeating the same song over and over, or playing the same artist twice in the same hour, are all gone. The boring parts decide whether you can actually broadcast. Royalty reports, listener stats per song, loudness that holds across a 1975 master and

2026-09-07 原文 →
AI 资讯

The 200 Came From a Rental

A pull request arrived after midnight with a README that claimed the API was already healthy. The coding agent had started a process, requested its own localhost, and treated a 200 as proof the service would run for everyone. That response was genuine inside a short-lived workspace, yet it said nothing about the laptop waiting on Monday. The reviewer stared at a green sentence printed on a host that nobody on the team could reopen. This pattern appears whenever a coding agent can execute commands, not merely suggest them, and reviewers misread the transcript. Developers treat the agent's shell as a preview of their laptop because both sessions speak bash and render similar fonts. The analogy fails like a hotel gym standing in for a home garage, familiar until one bolt size changes. Claims in the next sections are the ones that keep returning during review, then a fingerprint workflow that makes the rental visible. Myth: a bound port means the service is portable Agents love a bound port because it is a crisp success token that copies cleanly into a README. A process that answers on the sandbox does not encode libc, extra packages, file layout, or the user's group permissions. Health checks measure a moment on a host you do not retain, not a contract with the checkout that will survive merge. Treat a remote 200 as proof that some files ran once, then demand a second run on CI or a laptop. A useful correction is to refuse README claims that cannot be replayed from a clean clone of the branch. Ask the agent for the exact command sequence, the working directory, and the non-secret environment keys it exported during the run. Then execute that sequence locally with undocumented keys unset, unless they already exist in the team's dotenv template. If the local run dies on a missing header or a path the sandbox invented, the original green check was a rental. Myth: a free remote box is unofficial CI Teams under schedule pressure will point at agent logs the way they once po

2026-09-07 原文 →