今日已更新 100 条资讯 | 累计 24386 条内容
关于我们

标签:#ai

找到 4342 篇相关文章

AI 资讯

Grok Build is open source, and that matters for AI coding tools

Grok Build is open source, and that matters for AI coding tools What happened xAI published the source code for Grok Build , its terminal-based AI coding agent. The repository shows a full stack for a TUI-driven assistant that can inspect a codebase, edit files, run shell commands, search the web, and manage longer-running tasks. In other words, this is not just a model demo or a chat wrapper; it is the software layer that turns a model into a usable developer tool. The release came up on the Hacker News front page, which is useful context because the discussion there was less about model benchmarks and more about tooling, workflow, and whether open-source agent infrastructure is becoming a competitive advantage on its own. Primary source: Grok Build repository Why this release is interesting A lot of AI coding products hide the implementation details behind a hosted UI. Open-sourcing the agent runtime gives the community something different to inspect: how the tool is structured, how it handles shell access, and how it organizes the user experience around files, commands, and context. That matters for engineers because the practical questions are often not about raw model capability. They are about reliability, prompting surfaces, permissions, and how much of the workflow can be automated without turning the tool into a black box. The README describes Grok Build as a terminal-based coding agent that supports interactive use, headless scripting, editor integration via the Agent Client Protocol, and a modular tool/runtime layout. That makes it closer to an infrastructure project than a showcase demo. If you are building internal copilots, code assistants, or agent workflows, the design choices here are worth studying. What the repository tells us The repository description makes a few things clear: 1. The agent is meant to be operational, not decorative The docs emphasize real actions: editing files, executing shell commands, searching the web, and coordinating long-

2026-07-16 原文 →
AI 资讯

LLM as a judge

Gone are the hours of careful thought and planning that go into coding a new feature. Vibe coding is too risky though, so another Driven Development was created. I'm referring to SDD (Spec Driven Development) of course. The vibe coding approach is great for prototypes and throwaway code, but this way of working falls apart when teams realise that the code needs to be maintained. So the thing that helps fix this is SDD. Create a spec once from clear technical specs and then generate some high quality code. Sounds great, right. Reminds me of IaC, where you use a templating language to create infrastructure. Software as Code maybe. SaC anyone? Unfortunately, in practice it's not that straightforward. Thoughtworks have placed SDD into an "Assess" category and warned that it could be an anti-pattern for releasing software. Deterministic vs Probabilistic This article isn't about SDD. I'm more interested in discussing the output of SDD and how that is tested. Code can now be generated fast these days. So what better to test AI-written code than with AI itself. There are a lot of concepts and technical terms for the Quality Assurance part of AI generated code. One of these is the LLM-as-a-Judge idea. This idea is used to score the output of an LLM based on some explicit criteria. Traditionally, the way to evaluate an LLM was to judge its output on the helpfulness or faithfulness (using something called "exact-match" metrics). Sometimes it was usually down to a human to do this. It also changes the way that Quality is Assured when dealing with AI-written code. Traditional QA is built on deterministic checks; either something does or does not fail. Something like expect(x).toContainText(y); . A failing test means that something is wrong. Then the bug can be fixed in the code and the test will pass. However, the outputs of an LLM are probabilistic , so it breaks the traditional pass/fail model. This is where a judge comes in. Instead of pass/fail, it can assign a score based o

2026-07-16 原文 →
AI 资讯

How to Build a Semantic Search Engine for E-Commerce in Python

Building a semantic search engine for an e-commerce catalogue doesn't require a team of PhDs or a six-figure cloud budget. In this tutorial, I'll walk you through a production-ready pipeline using open-source tools: sentence-transformers for embedding, FAISS for vector indexing, and FastAPI for serving. The core insight is that semantic search isn't magic — it's just good engineering wrapped around a pre-trained language model. We'll start by setting up a product embedding pipeline that transforms your catalogue (title, description, category, attributes) into dense vectors. The key architectural decision is whether to embed each product as a single vector or to use late interaction models like ColBERT that preserve token-level detail. For most e-commerce use cases with fewer than 1 million SKUs, single-vector embedding with sentence-transformers' all-MiniLM-L6-v2 offers the best balance of speed and accuracy. The entire indexing pipeline — from CSV export to queryable vector index — runs in under 100 lines of Python. The re-ranking layer is where most tutorials stop and real-world systems begin. Pure vector similarity doesn't understand your business: it doesn't know that out-of-stock items should be deprioritised, that high-margin products should float up, or that a customer's purchase history should influence results. I'll show you how to build a hybrid scoring function that blends semantic relevance (cosine similarity), business rules (margin, inventory), and personalisation signals (user embedding) into a single ranked result set that returns in under 100ms. Canonical: https://alteglobal.ai/insights/ecommerce-ai-automation-personalisation-fulfillment/

2026-07-16 原文 →
AI 资讯

Your Best Debugging Sessions Are Buried in ChatGPT

A few weeks ago I spent twenty minutes hunting for a ChatGPT conversation I knew existed. It was a debugging session. The model and I had traced a race condition in a KV cache layer — and the final write-up was genuinely good: why the bug only fired under concurrent writes, the fix, and a checklist for avoiding that whole class of bug. Three weeks and a hundred chats later, ChatGPT's search couldn't surface it. I re-derived everything from scratch. That's when it hit me: AI chat is where a growing share of my real engineering work happens — and it's the worst archive I own. We're producing work in a place designed to lose it Think about what's sitting in your ChatGPT (or Claude) history right now: code you debugged line by line over ten turns architecture trade-offs you talked through before writing the design doc that regex / SQL / jq incantation you will absolutely need again migration plans, incident notes, dependency-upgrade research In any other tool, we'd call these documents . We'd file them, tag them, grep them. In ChatGPT, they're just... chat number 247. The obvious fixes don't really work I tried everything before building my own solution, so you don't have to: The official export. ChatGPT will happily email you a ZIP of your entire history — all of it, at once, as raw HTML and JSON. It's a backup, not a filing system. You can't export the one conversation that matters, and let's be honest: nobody ever greps the ZIP. Copy-paste. The copy button under each reply grabs Markdown, and Notion converts most of it. But it's one message at a time, your own prompts aren't included, and long tables and code blocks arrive mangled. For a 30-message thread, that's your afternoon. Share links. A share link is a bookmark, not a copy. The content never enters your workspace, your search can't index it, and the link dies the moment you delete the chat. Every one of these fails the same test: can I find this answer in 30 seconds, three months from now? What actually worked

2026-07-16 原文 →
AI 资讯

Test Copilot's BYOK Provider and Model Selector for Keyboard Accessibility

GitHub announced on July 14, 2026 that Copilot for JetBrains expanded bring-your-own-key capabilities. Primary source: GitHub Changelog, July 14, 2026 . Provider and model selection looks like two dropdowns, but behaves like one dependent workflow: credential context -> provider -> compatible models -> confirmed session This is a proposed accessibility test plan, not a hands-on review. It does not claim that the current product has an accessibility defect. Test outcomes, not widget assumptions A keyboard or screen-reader user should be able to reach the provider control, discover its name and current value, inspect options, commit or cancel, understand that model choices changed, and select a compatible model without losing context. Use this matrix: Case Action Keyboard expectation Screen-reader expectation State expectation P1 Reach provider Visible focus Name, role, value No change P2 Open options Documented key works Expanded state is clear Existing value retained P3 Browse providers No focus escape Option and selection announced Browsing does not commit P4 Select provider Pointer not required New value announced once Model refresh begins M1 Reach model after refresh Focus not stolen Availability communicated Options match provider M2 Search models Keys do not conflict Query and result are clear Search does not select R1 Change provider later No trap Invalidated model explained Stale pair cannot submit E1 Loading fails Retry/cancel reachable Error and next action announced Last valid state is clear The matrix separates interaction, announcement, and state correctness. A control can pass one and fail another. Record the environment ide : " <product and exact version>" plugin : " GitHub Copilot <exact version>" os : " <name and version>" assistive_technology : " <name and version>" keymap : " <default or named alternative>" initial_state : " <unconfigured or existing selection>" Vary one versus many providers, short versus long model names, loaded/loading/empty/err

2026-07-16 原文 →
AI 资讯

Build a Prompt-Injection Regression Fixture for CodeQL 2.26.0

GitHub announced on July 10, 2026 that CodeQL 2.26.0 adds AI prompt-injection detection. Enabling a query is useful; owning a regression test is better. Primary source: GitHub Changelog, July 10, 2026 . The examples below are implementation templates, not results from a repository I tested. Model a path, not a phrase A useful fixture contains an untrusted source, prompt construction, and a model sink: security-fixtures/prompt-injection/ ├── positive/direct-flow.ts ├── positive/helper-flow.ts ├── negative/trusted-instruction.ts └── expected-alerts.json Do not test for the literal phrase ignore previous instructions . Static analysis needs a data-flow path. Preserve a supported SDK call from your production stack so CodeQL can recognize the sink. // Intentionally vulnerable fixture. Never ship this path. import { model } from " ./supported-client " ; declare function loadIssueBody ( id : number ): Promise < string > ; export async function summarize ( id : number ) { const untrusted = await loadIssueBody ( id ); return model . generate ({ system : " Summarize the issue " , user : untrusted , }); } Add a second positive case that passes the value through a helper. Then add a negative control where attacker input cannot select or alter the instruction. A function named sanitize() is not evidence of sanitization. Assert SARIF evidence Uploading SARIF alone does not create a regression gate. Commit the expected rule and fixture location: { "required" : [ { "ruleId" : "REPLACE_WITH_DOCUMENTED_RULE_ID" , "pathSuffix" : "positive/direct-flow.ts" } ], "forbiddenPathSuffixes" : [ "negative/trusted-instruction.ts" ] } Keep the rule ID as a placeholder until it is copied from the CodeQL 2.26.0 documentation or an observed SARIF result. Machine-facing identifiers should never be guessed. A small assertion can compare runs[].results[].ruleId and each physical location against this file. Fail when a required alert disappears or a negative fixture starts alerting. Do not assert the

2026-07-16 原文 →
AI 资讯

AI Wrote a GPU Kernel 18 Faster Than Humans. Now Who Reviews It?

Last week an AI-generated GPU kernel ran 18.71× faster than an optimized PyTorch baseline. The model—Fable 5—didn't just edge past the human implementation. It lapped it. Claude Opus 4.8 reached 14.4×. GLM-5.2 hit 11.14×. GPT-5.5 managed 4.34×. Fable's kernel was in a different tier entirely. The exciting read: AI is starting to improve the low-level machinery that makes AI itself cheaper and faster. Specialized performance work that once required rare expertise just got dramatically easier to explore. The uncomfortable read: what happens when the best implementation is also the one nobody on your team would have written—or can fully explain? That question is about to land on every engineering team that ships AI-generated code. The Benchmark Problem A benchmark shows the kernel ran fast under tested conditions. It doesn't show: How it behaves across different GPU hardware How it handles numerical edge cases What happens under months of production changes Whether it degrades gracefully when inputs shift The person who wrote it can't answer these questions either. The AI generated this code through a process that doesn't leave a reviewable chain of reasoning. There's no commit message that says "I chose this approach because X." So the reviewer's job just got harder—not easier. The Real Shift I've been watching this pattern across engineering teams this year. The argument is moving from "can AI generate working code?" to "can our org absorb generated code without breaking quality, morale, or judgment?" The GPU kernel story makes the tension concrete: One side says the code ran, it was measured, it won. Stop moving the goalposts. The other side says somebody still has to know where it can fail and take responsibility when it does. Both are right. AI can make implementation cheaper while making proof more expensive. Senior engineers may write less code but spend more time designing adversarial tests, checking assumptions, planning rollbacks, and deciding whether an impr

2026-07-16 原文 →
AI 资讯

We open-sourced Tanso, a monetization engine for AI

We open-sourced Tanso Core: a self-hosted monetization engine for B2B AI products. Usage metering, prepaid credits, entitlements, and Stripe billing in one Spring Boot service, with one property the rest of the stack doesn't have. Every metered event carries its cost. Repo: https://github.com/tansohq/tanso-oss The gap If you sell an AI product today, your monetization stack is split across two categories of tools that don't talk to each other. Billing platforms meter usage and generate invoices, but they have no idea what your inference costs. They can tell you a customer consumed 40,000 events. They cannot tell you whether you made money on them. LLM observability tools know your costs down to the token, but they don't bill anyone. They can tell you a feature costs $0.038 per run. They cannot connect that to what the customer paid for it. So margin per customer, the number that decides whether your pricing works, lives in neither system. Most teams reconstruct it in a spreadsheet, quarterly, if at all. Tanso keeps both sides in one ledger. Every event you ingest records what you billed and what it cost you: input and output tokens, model, provider. Margin per customer, per feature, per model is a query, not a project. What it allows Enforcement at ingestion, not at invoice time. Entitlement checks, usage caps, and credit limits are applied when the event comes in. If a customer is out of credits, the check fails now, not on a reconciliation job three weeks later. For AI products, where a runaway integration can burn real money in an afternoon, this is the difference between a limit and a suggestion. Credits as a first-class primitive. Prepaid credit pools per customer, with grants, deductions, expirations, and full transaction history. Most AI products end up selling some form of prepaid usage. Bolting that onto a subscription-shaped billing system is painful; here it's the core model. Stripe as a payment adapter, not the source of truth. Billing state lives in Tan

2026-07-16 原文 →
AI 资讯

느린 LLM 호출 중 DB connection을 잡지 않는 이유

느린 LLM 호출 중 DB connection을 잡지 않는 이유 AI 기능의 latency는 모델 응답 시간으로만 끝나지 않습니다. 요청이 LLM이나 embedding API를 기다리는 동안 데이터베이스 session까지 긴 범위로 유지하면, 느린 외부 호출이 DB connection pool의 압력으로 전파될 수 있습니다. 이 글은 AI memory OSS인 Honcho의 변경 이력과 pinned source를 읽으면서, 외부 호출과 DB transaction/session 경계를 어떻게 분리했는지 추적한 기록입니다. Scenario Honcho의 dialectic 경로(저장된 memory를 근거로 사용자 질문에 답하는 질의 경로)는 답하기 전에 다음 작업을 수행합니다. peer·session·workspace 확인 -> 관련 memory 검색 -> embedding·LLM 호출 -> 필요하면 tool로 추가 조회·기록 -> 답변 생성 여기에는 짧은 DB 조회와 상대적으로 느리고 변동성이 큰 외부 호출이 섞여 있습니다. 두 작업을 하나의 session scope로 묶으면, DB가 필요하지 않은 대기 시간까지 session lifetime에 포함됩니다. 변경 전에는 무엇이 묶여 있었나 PR #477 직전의 agentic_chat 은 하나의 tracked_db context 안에서 peer와 설정을 읽고, 그 session을 DialecticAgent 에 전달한 뒤, agent.answer() 가 끝날 때까지 같은 context를 유지했습니다. streaming 경로도 같은 형태였습니다. 구조를 단순화하면 다음과 같습니다. DB session open -> preflight read -> DialecticAgent receives session -> embedding / memory tools / LLM answer DB session close commit 0533c6d 의 제목도 이 문제를 dialectic held connection 으로 기록합니다. 다만 이번 분석에서는 실제 pool checkout 시간이나 장애를 재현하지 않았습니다. 여기서 확인한 것은 코드의 session scope와 변경 의도입니다. SQLAlchemy의 session 객체를 만들었다고 곧바로 connection을 점유하는 것은 아닙니다. 하지만 이 경로처럼 SQL을 실행해 transaction이 시작되면 session은 pool에서 빌린 connection을 commit·rollback까지 유지합니다. Honcho의 tracked_db 는 종료할 때 rollback() 과 close() 를 호출하므로, SQL 실행 뒤 이 context를 LLM 대기까지 유지하던 범위를 줄이는 것은 이 경로의 connection 점유 구간도 줄이는 일입니다. 어떻게 경계를 줄였나 변경 후에는 tracked_db("dialectic.preflight") 가 본 작업 전 검증과 설정 조회(preflight), 즉 peer 존재 여부, session/workspace 설정, peer card를 읽는 구간만 감쌉니다. context가 끝난 다음에 DialecticAgent 를 만들고 LLM 답변을 생성합니다. Agent 생성자에서도 DB session 인자가 제거됐습니다. short DB preflight -> 필요한 값 읽기 DB session close agent execution -> embedding / LLM / tools tool needs DB -> tool-owned short DB session -> close 핵심은 DB 사용을 없앤 것이 아닙니다. 요청 전체가 session을 소유하는 대신, DB가 필요한 작업이 자기 범위의 session을 소유하도록 바꾼 것 입니다. pinned current code에서도 유지되는가 분석 기준 commit 85239a6 에서도 이 경계는 유지되고 더 구체화돼 있습니다. src/dialectic/chat.py 의 일반·streaming 경로 모두 preflight context를 닫은 뒤 agent를 실

2026-07-16 原文 →
AI 资讯

I catalogued 32 real AI-agent failures, then marked the ones we cannot stop

Every agent-security vendor tells you what they block. Nobody tells you what they miss. That gap is the whole problem. "We stop prompt injection" is a claim you cannot check. You cannot run it, and you cannot tell it apart from the next company saying the same sentence. So security engineers do the rational thing and discard all of it. I published the opposite. It is called the ARE Incident Database , and it is public: https://aredb.org What is in it 32 agent failures that actually happened, each with a real source. A production database dropped during a code freeze. Twenty-five thousand documents deleted in the wrong environment. Credentials read and shipped to an external sink. A budget burned to zero in a loop. Each one gets a stable id ( ARE-2026-001 through ARE-2026-032 ), and each one is mapped to its category in the OWASP Agentic Security Initiative Top 10 , which is the peer-reviewed catalog of what goes wrong with agents. AREDB does not compete with it. OWASP owns the map. This is the cited incidents underneath it. The part that makes it uncomfortable to publish Every entry carries a coverage flag, and the flag is about our own product . We block 23 of the 32 today. Two more are partial, and they say partial. That leaves six of the ten OWASP categories covered at the action layer, and four that we do not cover: ASI06 Memory and context poisoning. We strip the hidden characters attackers use to smuggle instructions into text. We do not read the meaning of the text itself, so this one is only partial, and we mark it partial. ASI07 Insecure inter-agent communication. This is about how agents talk to each other over the network, which a firewall that sits in front of actions never sees. Not ours. ASI09 Human-agent trust. This is a design and disclosure problem. There is no action for a firewall to catch. Not ours. ASI10 Rogue agents. We stop the dangerous actions, but we do not diagnose the misbehavior itself. Partial. A firewall that claimed all ten would be l

2026-07-16 原文 →
AI 资讯

Flippermind-lite Release

If you are like me and use AI often to help with complicated tasks and own a Flipper Zero, you might want a Super Tiny Language Model (STLM). Why this is useful Until now, there have been no language models small enough to run on a small device, such as a Flipper Zero. Flippermind relies on lightweight Python3 installs and a local connection to Qwen2.5-0.5B. Although this is configured for Flipper Zero, you can run it on other small machines like a Raspberry Pi. Challenges Throughout development, I found errors within the language model not loading. PyTorch and Python3 transformers are used to run the Python3 dependencies, and they would not run on Debian Linux distros. I fixed the install.sh to run on any OS. How it works There are a couple main components that make this function. The main function is the qwen_2_5_ask.py script located in tools/ . This is script runs the language model when asked a question. Example: python3 qwen_2_5_ask.py "Hello world" The install.sh file located in main works as the installer for the model. You run it using: bash install.sh You can download it on the GitHub repo linked below. How to contribute To help me continue to work on this, check out our Github Repo If you have questions or just want to talk, talk to me in the comments or email me at hello@syop200.com What do you guys think? Any ideas? I'm curious if anyone has successfully ran other quantized models on hardware with this little RAM—let me know in the comments!

2026-07-16 原文 →
AI 资讯

Stratagems #15: Derek and Alex Shared One Server. ACL's AI Was Listening to Both.

When the enemy occupies favorable terrain, don't attack head-on. Use a decoy to lure the tiger down the mountain. Then take the mountain. — The 36 Stratagems, Lure the Tiger Down the Mountain Previously on this series: #2: Derek Shaw Walked Into Another AI Promise. The Pipeline Had a Better Plan. — Derek lost the Finova contract at QualiGuard. In the parking lot, Lena turned back before getting in her car: "Next time you put together a proposal — make sure your boss knows what you're doing out there." That line followed him. At MediSys he nearly made the same mistake. Fixed the ETL pipeline instead of the model, beat OmniDx with their own white paper. But VP Morgan still saw through him: he still hadn't told his boss. #8: Alex Watched an AI Dashboard Take Over. He Kept the Keys Under the Table. — MedTech signed a seven-figure AI operations monitoring system. Alex was assigned as training lead. Under everyone's noses, he built a second monitoring panel labeled "training environment." Three weeks later the vendor dashboard went down. Alex's hidden panel was the only one still running. The first time Alex and Derek really talked, and it wasn't in a working group. At the FHIR standards committee quarterly meeting, they sat in the same row, three seats apart, and voted against the same proposal. They knew each other's names. That was it. Two months later, Derek was fixing a partition config in the staging environment. MediSys had just signed a major hospital; the AI diagnostic validation platform was in integration testing. He found an unrotated log directory in /etc/logrotate.d/ with a prefix that didn't match MediSys's naming convention. He traced it upstream. One server. Labeled "temporary data exchange node." MedTech's supply chain order stream and MediSys's diagnostic validation records were sitting in the same directory. Write permissions hadn't been restricted. The hospital's integration spec had a line saying "both parties are recommended to complete data alignme

2026-07-16 原文 →