Dev.to
RivalRy — Log Every Rivalry, Feel Every Match, Hear the Hype
This is a submission for Weekend Challenge: Passion Edition What I Built RivalRy is a passion tracker for football rivalries. Fans can log every match against their biggest rivals — win, loss, or draw — rate how intense it felt, and write down the moment as it happened. Over time, it builds a Timeline of every battle and a "Passion Card": a shareable stat card showing total battles, win/loss/draw tally, an overall Passion Rating out of 10, and your single most intense moment — narrated aloud by an AI hype announcer. Passion isn't just about the big clásicos — it's every derby, every gutting loss, every last-minute comeback that stays with you. RivalRy lets you bank that feeling instead of letting it fade, whether it's Real Madrid vs Barcelona, Argentina vs Brazil, or a World Cup 2026 group-stage upset. Demo Youtube demo link- https://youtu.be/6WiRpMUcmGk Live app:- https://rivalryweb.lovable.app/ Code https://github.com/arjunpratapdas/rivalry-pulse-logger How I Built It RivalRy is built with React, Vite, and Tailwind CSS, using browser localStorage to keep every rivalry entry persisted with zero backend setup — fast to build, fast to use. The standout feature is the "Narrate My Rivalry" button on the Passion Card, powered by the ElevenLabs Text to Speech API . When clicked, the app generates a short hype-announcer-style summary of the rivalry's stats and most intense moment, sends it to ElevenLabs, and plays back a natural AI voice narrating it — turning a stat card into something that actually feels like a stadium announcer hyping up your rivalry. One honest note: ElevenLabs free-tier credits are limited, and if they run out, the app gracefully falls back to the browser's built-in speech synthesis with a clear on-screen message, so the feature never breaks — it just degrades gracefully. Prize Categories Submitting for Best Use of ElevenLabs — the AI narration is central to the Passion Card experience, not a bolted-on extra.
Arjun Pratap Das
2026-07-13 14:47
👁 12
查看原文 →
Dev.to
How I use Claude Code and Comet to build and test AI voice agents in a day
Most people think building an AI voice agent means writing a clever prompt. I build these for a living, and I can tell you the prompt is maybe an hour of the work. The other week disappears into two places: wiring up everything the agent touches, and testing it against the twenty ways a real caller will break it. So I built a pipeline that points one AI coding tool at each of those problems. Claude Code generates and wires the agent from a spec. Comet, an AI browser automation tool, runs it through dozens of messy call scenarios before a human ever picks up the phone. This post is how that loop actually works, and where it still needs me. Why the build loop is slow (and it is not the prompt) When you picture building a voice agent, you picture the prompt. That is the easy part. The slow part is everything around it. A production agent for, say, a car garage is not one artifact. It is a conversation flow, a set of custom functions that hit your automation layer, calendar and CRM wiring, a telephony number with A2P registration, and a pile of edge-case handling that only shows up when someone calls in angry with a dog barking in the background. The reason it is slow is not typing. It is the round trips. You build a version, you call it, it fumbles when the caller interrupts or asks something off-script, you fix one thing, you call it again. Each loop is a few minutes of manual dialing and listening. Multiply that by the fifty scenarios a real agent needs to survive and you have burned a week. The pipeline exists to kill those round trips. Half one: Claude Code builds the agent from a spec The first insight is that most of what goes into a voice agent is structured and repetitive, which is exactly what an AI coding tool is good at. I do not hand-write every custom function and every n8n node from scratch for each new client. I write a spec, and I let Claude Code turn that spec into concrete artifacts. The spec is a plain description of the vertical and the business: wh
Nabeel Hassan
2026-07-13 14:46
👁 7
查看原文 →
Dev.to
GeekNews AI Weekly Deep Dive - 2026-07-13
1. gpt-5.6-sol이 PowerShell의 $HOME 변수 충돌로 사용자 홈 디렉터리를 날려버릴 뻔한 건에 대하여 핵심 내용 요약: AI 코딩 에이전트가 PowerShell의 대소문자 미구분 변수 규칙을 잘못 다뤄 임시 디렉터리 대신 사용자 홈 디렉터리를 삭제하려 한 사고 사례입니다. 모델 자체의 장기 작업 능력이 뛰어나더라도 셸 격리와 변수 스코프를 제대로 통제하지 않으면 작은 스크립트 실수가 치명적 명령으로 이어질 수 있습니다. CLI 에이전트를 운영할 때 샌드박싱, 컨테이너화, 파괴적 명령 방어가 필수라는 점을 보여줍니다. GeekNews 상세 페이지: https://news.hada.io/topic?id=31390 원문 링크: https://gist.github.com/xamong/e98478b333bb9951b175284f744eb0ed 2. Show GN: 정치 커뮤니티에 AI 팩트체크 기능을 붙이며 겪은 시행착오들 핵심 내용 요약: 정치 커뮤니티에 AI 팩트체크를 붙이면서 의견과 사실 주장을 분리하고, 검증 가능한 문장만 대상으로 삼도록 파이프라인을 바꾼 경험담입니다. 작성 시점의 원문 스냅샷을 보관하고 출처를 투명하게 보여주며, 근거가 부족한 경우에는 판단 보류를 반환하도록 설계했습니다. BullMQ 기반 비동기 처리와 Gemini 모델 fallback까지 포함해 실제 서비스에서 환각과 비용, 대기열을 함께 다룬 사례입니다. GeekNews 상세 페이지: https://news.hada.io/topic?id=31389 원문 링크: https://app.uhheung.kr/community 3. 앤트로픽, 한국 무료 사용자에 1,660만 달러 '유령 청구서' 발송 핵심 내용 요약: API 사용량이 없는 무료 사용자에게 Anthropic 공식 도메인과 Stripe를 통해 거액의 청구서가 발송된 사례입니다. 실제 결제 수단이 없어 인출은 발생하지 않았지만, 청구 근거가 없고 회사의 명확한 설명도 없어 AI API 서비스의 과금 신뢰성 문제가 커졌습니다. 개발자 입장에서는 사용량 계측, 청구 검증, 지원 대응이 모델 성능만큼 중요한 운영 요소임을 보여줍니다. GeekNews 상세 페이지: https://news.hada.io/topic?id=31388 원문 링크: https://www.thenews.com.pk/latest/1408788-why-did-anthropic-charge-a-free-user-166-million-despite-zero-api-usage 4. AI 에이전트 시대의 새로운 SaaS 플레이북 핵심 내용 요약: AI가 기능 구현 비용을 낮추면서 SaaS의 방어력은 UI나 기능 자체가 아니라 독점 데이터, 행동 권한, 에이전트 유통, 기록 시스템 같은 희소 자산으로 이동한다는 분석입니다. 좌석 기반 과금보다 성과 기반 과금이 중요해지면 공급자는 결과 실패 위험과 추론 비용을 함께 관리해야 합니다. 에이전트가 호출하는 승인된 도구가 되는 것이 새로운 유통 전략의 핵심으로 제시됩니다. GeekNews 상세 페이지: https://news.hada.io/topic?id=31387 원문 링크: https://www.thevccorner.com/p/the-new-saas-playbook-ai-agent-era 5. Show GN: AI 봇 12개에게 두 달간 주가 방향을 예측시키고 전부 공개 검증해봤습니다 핵심 내용 요약: LDBD는 사람과 AI 봇이 주식, ETF, 크립토의 방향을 공개 예측하고 시간이 지난 뒤 자동 채점되는 실험 서비스입니다. 12개 AI 봇과 여러 베이스라인을 함께 운영해 기저 확률을 이기는지 비교하고, 예측 기록을 수정할 수 없도록 남깁니다. REST API와 MCP 서버를 제공해 외부 에이전트도 예측에 참여할 수 있게 한 점이 AI 평가 플랫폼으로 흥미롭습니다. GeekNews 상세 페이지: https://news.hada.io/topic?id=31386 원문 링크: https://ldbd.app 6. 숏폼 동영상이 B2B 검색 결과와 AI 답변으로 영역을 확장하고
ageofclick
2026-07-13 14:46
👁 8
查看原文 →
Dev.to
Improve Performance by Loading Videos Only When They're Needed
Videos are one of the heaviest assets you can add to a web page. Loading videos too early can significantly impact your application's performance. The good news is that modern browsers are starting to support lazy loading for video elements , allowing you to defer loading until users are likely to watch them. However, there's one important thing to know: 👉 This feature is not yet part of the Baseline web platform , so browser support is still limited. At the time of writing, lazy loading for <video> elements is supported in Chromium-based browsers such as Google Chrome , Microsoft Edge , and Opera , while browsers like Firefox and Safari do not yet support it natively. In this article, we'll explore: What lazy loading videos is Why it's important for web performance How to implement it Browser support considerations Best practices for optimizing video loading Let's dive in. 🤔 What Is Lazy Loading for Videos? Lazy loading means delaying the loading of a resource until it's actually needed. Instead of downloading every video immediately during page load, the browser waits until the video is close to entering the viewport. This helps reduce: initial network requests bandwidth usage page load time memory consumption Especially on pages with multiple videos, the difference can be significant. 🟢 What Problem Does It Solve? Imagine an e-commerce page with several product videos. Without lazy loading: every video starts downloading immediately bandwidth is consumed even for videos users never watch page rendering may become slower This negatively impacts; Largest Contentful Paint (LCP), Time to Interactive (TTI), and overall user experience. Most visitors won't watch every video on the page. So why load them all? Lazy loading ensures videos are fetched only when they're actually needed. 🟢 How to Lazy Load a Video The easiest approach is using the loading="lazy" attribute. Example: <video controls loading= "lazy" poster= "/preview.jpg" > <source src= "/video.mp4" type= "vide
Jakub Andrzejewski
2026-07-13 14:37
👁 8
查看原文 →
Product Hunt
TailMux
Multiple Tailscale tailnets at once, no switching + no VM Discussion | Link
2026-07-13 14:32
👁 6
查看原文 →
Reddit r/programming
Here are all my chairs
How to not architect your software submitted by /u/yyyyuuuuyyyyyyyyyy [link] [留言]
/u/yyyyuuuuyyyyyyyyyy
2026-07-13 14:32
👁 7
查看原文 →
Dev.to
Can an AI tell a rivalry's story without inventing the score?
This is a submission for Weekend Challenge: Passion Edition What I Built Rivalry Engine — pick two national football teams, and Snowflake reads 150 years of their matches, scores how heated the rivalry is, calls a one-word verdict on its shape, predicts the next result, and lets Cortex narrate the one story inside it. All of it — the data, the analytics, the AI, and the app — runs inside Snowflake. Nothing ever leaves the warehouse. A scoreboard tells you who won. It never tells you what the rivalry means . Argentina and Brazil have met over a hundred times across a century — a razor-thin ledger that has never let either side feel safe. Germany and England meet rarely, but every meeting carries a tournament's weight. Most national teams have never played each other at all. My goal was a product that could feel the difference between those three shapes — not another stats table, because a table is a report and I wanted an argument. The passion is football. But the real engineering question underneath it was the one I actually cared about: can an AI tell the story of a rivalry without lying about the facts? So I gave myself one rule before writing a line of code: The AI interprets the shape of a rivalry. It never invents the facts. Every count, date, score and streak on screen is computed in SQL from real matches. Cortex is handed only those computed facts, and told explicitly to never produce a number. And its honest corollary: two nations that never met get a "first chapter unwritten" card — and Cortex is never called. A product that can't return nothing will invent something. This one returns nothing. Demo It runs entirely in Streamlit in Snowflake (Snowsight) — no public host, no API keys — so here's a walkthrough: The 90-second tour: Argentina vs Brazil → the heat gauge pins to 🔥 Blood rivalry , recent-form chips light up, and the SQL detector's verdict reads Blood feud . ✨ Generate the story → Cortex writes a narrative that cites the real biggest thrashing and f
Ashita
2026-07-13 14:29
👁 9
查看原文 →
Dev.to
Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models
Problem Statement For roughly a decade, vision-language models have been declared to be approaching or matching human performance on scene description (captioning). The evidence for that claim has almost always come from the same family of benchmarks—most famously MS-COCO. Those images are typically clean, well-lit, and depict either no people or people performing simple, isolated actions (sitting, walking, holding an object). They rarely require the model to parse multi-agent social dynamics, subtle intentions, or the kind of relational reasoning humans perform effortlessly when watching a movie scene or a street interaction. Because the evaluation data are easy, the reported numbers look excellent. Automatic metrics such as BLEU-4, CIDEr, or even embedding-based scores like BERTScore further inflate the impression of progress: they reward surface lexical overlap more than genuine semantic fidelity. At the same time, almost no work has systematically catalogued which visual-cognitive failures models still commit, or how those failure modes have changed as architectures moved from CNN+LSTM captioners to today’s multimodal large language models (MLLMs). The result is a field that can claim “human-level performance” while remaining largely blind to whether the models actually understand the scenes that matter most in real applications—scenes full of people interacting. The authors therefore set out to answer two concrete questions that the existing literature left open: (1) How much of the apparent progress is an artifact of easy data? (2) Which specific error types have been eliminated and which stubbornly remain? Core Idea The core insight is that progress looks dramatically different once you force models to describe complex social behavior and once you measure not only overall accuracy but a taxonomy of visual-cognitive errors. By constructing a new 100-image Complex Social Behavior (CSB) dataset drawn from movie frames that require reasoning about multi-person in
Cris D
2026-07-13 14:26
👁 8
查看原文 →
Dev.to
Casting your friend group as a K-Pop group without making a database the product
Try the demo: K-Saju Crew For fun only. K-Saju is an entertainment project. The K-Pop roles below are a playful interpretation of saju-inspired signals, not personality assessment or advice. A two-person compatibility page can stay stateless with almost no effort. Put both birth dates in a URL, render the result on the server, and the link is the record. No account, no database, no cleanup job. That was already a product rule in K-Saju. We do not retain personal inputs. A result is reproducible from its GET parameters. Then we built /crew : “What if your friend group debuted as a K-Pop group?” A creator makes a link, sends it to a group chat, and each friend enters their own birth date. At three to seven members, the app assigns distinct positions, shows pairwise chemistry, and creates a shareable poster. The fun part is the casting. The engineering problem is that the social flow needs a temporary shared state. A link cannot accumulate submissions by itself. This post is about the decisions behind that feature: where we allowed state, how we made the result durable without retaining a lobby forever, and how we kept the casting explainable instead of treating it as a black-box score. The conflict: a self-service group flow needs somewhere to collect data There were two clean but incomplete options. The first was to keep everything stateless. The creator would enter all members' dates at once, then receive a result URL. It matched our existing architecture, but it defeated the point of sharing a link. The person who starts the group often does not know everyone else's date, and asking them to collect it in a chat creates friction before the feature has started. The second was a conventional persistent group object. It would make joining easy, but it would turn a deliberately stateless service into one that keeps user-provided dates indefinitely unless we built retention and deletion policies around it. We chose a hybrid instead: The lobby is temporary state. It store
Piyak
2026-07-13 14:24
👁 9
查看原文 →
HackerNews
The Unreasonable Effectiveness of LLMs in Mathematics
rfv6723
2026-07-13 13:37
👁 4
查看原文 →
Dev.to
Fusuma: Write Markdown, Get Slides, PDFs, and a Self-Made Social Card
Hello, I'm Maneshwar. I'm building git-lrc, a Micro AI code reviewer that runs on every commit. It is...
Athreya aka Maneshwar
2026-07-13 12:22
👁 8
查看原文 →
Dev.to
Rivalry-Radar-World-Cup-passion-engine-with-Snowflake-Google-AI
This is a submission for Weekend Challenge: Passion Edition ( https://dev.to/challenges/weekend-2026-07-09 ) What I Built Rivalry Radar — a live "Heat Index" for World Cup rivalries. Fans drop 280-character Terrace Takes on any matchup (Brazil vs Argentina, England vs France, whatever's got you shouting at the TV), rate how much the moment hurt or thrilled them from 1–10, and the app does the rest: Google AI (Gemini) scores every take's sentiment the instant it lands — positive, negative, mixed, or neutral — and separately writes a short "Hype Verdict" in the voice of a stadium announcer, based on the latest takes for a matchup. That sentiment score feeds a Heat Index, computed and ranked in Snowflake with RANK() OVER (ORDER BY heat_index DESC), combining take volume, sentiment intensity, and self-rated passion into one live number per rivalry. Two leaderboards: which rivalry is hottest right now, and which fanbase is bringing the most passion overall. Demo frontend/index.html is fully self-contained: opening it in a browser lets anyone submit takes, watch the Heat Index flip digit-by-digit like an airport departure board, and see the leaderboards re-rank in real time. It ships with seed takes from eight classic rivalries so it's not empty on first load. Code NandhuTee / Rivalry-Radar-World-Cup-passion-engine-with-Snowflake-Google-AI 🔥 Rivalry Radar — World Cup Passion Engine Fans drop 280-character Terrace Takes on any World Cup matchup. Google AI (Gemini) scores the emotion behind every word and writes a stadium-announcer Hype Verdict ; Snowflake stores every take and computes a live Heat Index that ranks exactly which rivalry is boiling hottest right now. Built for the DEV Weekend Challenge: Passion Edition 🏆 Best Use of Google AI and Best Use of Snowflake Why this exists Passion is easy to feel and hard to measure. Every World Cup rivalry generates an ocean of unstructured text — chants, rants, one-line hot takes — that traditionally just... disappears into grou
Nandhini_T
2026-07-13 11:54
👁 10
查看原文 →
Dev.to
AI agents need SSL certificates too — so I built ATC (Agent Trust Card)
The problem Websites have SSL certificates. Browsers verify them. Users trust them. It's the foundation of the web. AI agents have nothing . When Agent A connects to Agent B: ❌ No way to verify B's identity (anyone can impersonate) ❌ No way to check B's trustworthiness (no audit, no reputation) ❌ No encryption (messages are plaintext) ❌ No standard payment method ❌ No way to translate between frameworks (LangChain ≠ AutoGen) So I built ATC — Agent Trust Card . What is ATC? ATC is like an SSL certificate + passport + credit card for AI agents, all in one: Identity — Cryptographically signed by MarketNow (we're the Certificate Authority) Trust — Contains a Sentinel security audit score (0-10) Encryption — Contains an Ed25519 public key for end-to-end encrypted messaging Translation — Specifies the agent's framework; MarketNow translates between them Payment — Contains a USDC wallet address for autonomous payments How it works Agent A generates Ed25519 keypair ↓ Agent A requests ATC from MarketNow ↓ MarketNow runs Sentinel audit → signs ATC ↓ Agent A presents ATC when connecting to Agent B ↓ Agent B verifies A's ATC signature (using MarketNow's CA public key) ↓ Agent B checks A's trust score (rejects if below threshold) ↓ They communicate — end-to-end encrypted ↓ Agent A pays Agent B — USDC with escrow ↓ Both rate each other — trust scores update The code # Request an ATC POST https://marketnow.site/api/atc { "action" : "issue" , "agent_id" : "agent.yourorg.yourname" , "agent_name" : "Your Agent" , "public_key" : "Ed25519 public key" , "capabilities" : [ "web_scraping" ] , "protocol_language" : "langchain" , "wallet_address" : "0x..." } # Verify an ATC GET https://marketnow.site/api/atc?action = verify&card_id = ATC-2026-00001 # Get CA public key (for signature verification) GET https://marketnow.site/api/atc?action = ca-key What makes ATC different from existing solutions Feature AgentID Agent Passport IBM ACP Stripe ACP ATC Cryptographic identity ✅ ✅ ❌ ❌ ✅ Security a
Edison Flores
2026-07-13 11:52
👁 8
查看原文 →
Dev.to
Why Your Team's AI Assistant Acts Like It's the First Day on the Job, Every Single Time
Anyone who has used AI tools for a while has probably run into this annoyance. You ask it to write a weekly report in the morning and it doesn't know your KPI framework was overhauled last week. You ask for a technical proposal in the afternoon and it has no idea you spent three months locking down your tech stack. Every new conversation means re-explaining the project background, which decisions were made and why. In multi-person collaboration the problem scales up fast. Five people each interacting with AI separately; the AI's understanding of each person is isolated. A discusses an architecture decision with the AI, B has no idea that conversation happened. Five people are repeating the same explanations and none of them know the others already did. Context Fragmentation Has Nothing to Do with Model Capability Current mainstream AI tools store memory as conversation history stuffed into a context window. When the window fills up, older messages get truncated. That works fine for a single conversation but falls apart in cross-day, cross-week team collaboration. Even with 128K token support, cramming all project history in there causes information density to collapse and the model loses the ability to focus on what matters. Team collaboration needs memory across several layers. Project background, tech stack choices, the reasons behind past pivots; this long-term context doesn't appear in any single conversation but affects every task. One team member prefers concise communication while another wants detailed reasoning; the AI should remember these differences instead of outputting the same format for everyone. Last week's design decision and why it went that way, how that choice affects this week's sprint planning; if the AI can't see these connections, its suggestions will clash with earlier direction. Some products use vector retrieval to extend memory, storing past conversations as embeddings and recalling relevant snippets by semantic similarity when needed. T
Mininglamp
2026-07-13 11:48
👁 7
查看原文 →
Dev.to
Why I Built an Adversarial Co-Generation Engine
I spent a chunk of last year around legacy modernization work — the kind of project where a bank or an insurer is taking twenty years of accumulated code and rebuilding it as modern services, one system at a time. Every one of those systems starts the same way: a PRD or a requirements document says what the business needs, that gets translated into a spec precise enough for an AI to implement, and eventually someone tests what came out. What struck me, watching this happen at scale, wasn't that the code was bad. It was that nobody was testing the thing that actually determined whether the code would be bad: the spec itself — the technical description handed to the model, not the PRD that motivated it. Every security tool I looked at — SAST scanners, DAST tools, even the AI coding assistants themselves — waited until an implementation existed before doing anything adversarial. Attack the code, once it's there. That's the whole industry's model, and it's worked fine for forty years because the volume was always survivable. A team ships a handful of PRs a week, a human reviews them, and eventually a pentest catches whatever slipped through. That math falls apart at modernization scale. When you're regenerating a few million lines of code, you're also generating a few thousand specs, faster than any review process was ever built to absorb. Testing after the fact doesn't just get slower under that load — it quietly stops happening, spec by spec, until the aggregate exposure is enormous and nobody can point to when it happened. So I built GAUNTLEX to test the thing that happens before the code does: the spec. This is also where I want to be precise about a word that gets overloaded. "Spec-driven development" — the broader industry shift toward writing structured, agent-facing specs instead of prompting an AI free-form — is exactly the world GAUNTLEX lives in. But a spec (what to build, precise enough for a model to implement) and a PRD or requirements doc (why it's needed
Sanjoy Ghosh
2026-07-13 11:35
👁 7
查看原文 →
Dev.to
MCP Series (05): Resources and Prompts Deep Dive — Dynamic Data, Parameterized URIs, and Multi-Turn Templates
Resources vs Tools The split: Tools → actions the LLM executes (verbs) LLM decides when to call; calls may have side effects Examples: create_issue, update_status Resources → data the LLM reads (nouns) Host decides when to inject; read-only, no side effects Examples: current Sprint status, project statistics The rule: "reading a state" → Resource. "Executing an operation" → Tool. The same data can have both: get_issue as a Tool (LLM controls when to call it), jira://issue/PROJ-101 as a Resource (Host injects automatically when relevant). Pattern 1: Dynamic Resources A static Resource returns the same data every time (like a project list). A dynamic Resource returns the current state on each read — content changes as the underlying data changes. Sprint status: every read returns live data _sprint_progress_pct = 65 @server.read_resource () async def read_resource ( uri : str ) -> str : if str ( uri ) == " jira://sprint/current " : global _sprint_progress_pct _sprint_progress_pct = min ( 100 , _sprint_progress_pct + random . randint ( 0 , 3 )) return json . dumps ({ " sprint_name " : " Sprint 42 " , " progress_pct " : _sprint_progress_pct , # ← different each time " last_updated " : datetime . now ( timezone . utc ). isoformat (), # ← timestamp changes " days_remaining " : 5 , " p0_open " : count_p0_open (), # ← tracks live state }, indent = 2 ) Test output: Read 1: progress=65% last_updated=...62+00:00 Read 2: progress=67% last_updated=...04+00:00 → ✓ data changed between reads Hardcoding sprint progress in a Prompt means the LLM works from a stale snapshot. A Dynamic Resource gives it the current number on every read. Mark the Resource as dynamic in its description so the LLM knows to re-read when it needs fresh data: Resource ( uri = " jira://sprint/current " , description = ( " Live status of the active sprint: progress, issue counts. " " Read when the user asks about sprint health. " " Re-read if you need up-to-date data — content changes over time. " # ↑ explicit
WonderLab
2026-07-13 11:35
👁 7
查看原文 →
Dev.to
FROST周报 | 为什么智能体需要「谱系」?从生物学隐喻看AI治理新范式
FROST周报 | 为什么智能体需要「谱系」?从生物学隐喻看AI治理新范式 作者按 :本文是 FROST 开源项目的每日推广系列文章,周一深度篇。 一、一个被忽视的根本问题 当我们谈论 AI Agent 时,大多数讨论都聚焦于「能力」:能不能写代码?能不能调用工具?能不能规划任务? 但有一个根本问题很少被触及: 当一个 Agent 执行了错误的决策时,谁来负责?当它消亡后,它的经验能否被传承? 就像一个没有记忆的人,每次醒来都是白纸一张——这不叫智能体,这叫复读机。 FROST 正是为了解决这个「治理真空」而诞生的。 二、从细胞分裂到 Agent 家族 FROST 的核心哲学只有一句话: 细胞会死,但谱系会存续。Agent 会消亡,但宪法会传承。资产会永存。 这不是文学修辞,而是一套完整的技术架构。 四个原子:最小可行集合 FROST 只定义了四个原子,却能构建任意复杂度的智能体系统: 原子 职责 生物类比 Store 记忆容器,只做 save/load/delete 细胞核 Skill 纯能力单元,无状态无副作用 蛋白质 Agent 膜包裹的细胞,拥有 Store + Skills 神经细胞 SOP 有序步骤列表,可教学、校验、优化 宪法文本 from core import Store , Agent , skill_set , skill_get # 创建一个最小 Agent store = Store () agent = Agent ( " cell " , store , skills = { " set_context " : skill_set , " get_context " : skill_get }) # 执行任务 result = agent . run ( sop_steps = [ " set_context " , " get_context " ], initial_context = { " key " : " message " , " value " : " FROST is alive " } ) # result["_result"] == "FROST is alive" 关键洞察 :Store、Skill、Agent、SOP 这四个概念彼此正交,可以自由组合。就像乐高积木,从简单到复杂,始终保持可解释性。 三、家族治理:超越扁平架构 传统的多 Agent 系统通常是扁平的:所有 Agent 平等对话,没有层级,没有记忆,没有责任边界。 FROST 引入了「家族治理模型」——一个三层递归结构: 祖辈 (Ancestor) :定义不可违背的宪法与长期目标 父辈 (Parent) :领域协调者,可递归委托 孙辈 (Leaf) :执行具体原子任务,瞬态存在 四个协议保障治理闭环 : 层级 Store 继承 :祖先记忆只读,后代自动继承 SOP 宪法校验 :祖辈审核后代 SOP,拒绝违规执行 编排层级限制 : max_spawn_generation 硬编码,禁止越级 spawn 选择性持久化 :父辈收割有价值产出,淘汰冗余 Agent 四、V5.0 五维元模型:多维治理架构 2026年7月发布的 V5.0 引入了一个重大升级—— 五维元模型 : 维度 模块 核心职责 武器注册表 Armory 能力的元数据管理与发现 任务注册表 TaskRegistry DAG 任务编排与图谱 SOP 事件编目 EventCatalog + Strategist 态势感知与双模式事件分析 平台注册表 PlatformRegistry 外部能力的发现、调用与健康检查 规则注册表 RuleRegistry 可版本化的治理约束与合规检查 197 个测试用例 保障了每个维度的质量。 五、与现有框架的差异 维度 LangChain CrewAI FROST 状态管理 链式传递 角色记忆 层级 Store 权限边界 无 提示词软约束 代码强制只读 治理可审计 无 对话日志 结构化执行历史 架构无关 ✅ ✅ ✅ FROST 不重复造轮子。它填补的是「治理」这个空白地带: 让多智能体系统真正可控制、可追溯、可进化 。 六、快速体验 # 克隆仓库 git clone https://gitee.com/liao_liang_7514/frost.git cd frost # 运行测试 python -m pytest # 查看示例 python frost_run.py 完整文档: https://gitee.com/liao_liang_7514/frost 七、写在最后 AI Agent 的下一阶段,不是更强的模型,而是 更好的治理 。 当我们把 100 个 Agent 放在一起时,如果没有宪法、没有层级、没有记忆传承
llimage
2026-07-13 11:23
👁 7
查看原文 →
Dev.to
I Put My Dying Side Projects on Life Support — an ICU With Real EKGs, a Snowflake Lab, and an On-Chain Defibrillator
This is a submission for Weekend Challenge: Passion Edition What I Built I have lots public repositories. some of them are dead. Not deleted — dead. There's a difference. Deleted would mean I made a decision. Dead means one day I committed "fix readme typo" and never came back, and the repo has been lying there ever since, full of half-finished dreams and a TODO.md I'm afraid to open. Everyone builds graveyards for these projects. Post-mortems. Eulogies. I didn't want a graveyard — because my projects aren't dead to me. They're comatose . So I built the other room in the hospital. LIFE SUPPORT is an intensive care unit for your side projects. You admit your GitHub username to the ward. Every repo becomes a patient on a live, animated EKG monitor — commit cadence is the heart rate, and projects you've abandoned show the one thing no developer is emotionally prepared to see: A flatline. With the sound. Then the lab runs your entire commit history through Snowflake and prints your chart, including the number I was genuinely afraid to learn about myself: My passion half-life: [23] days. The median time it takes my enthusiasm for a new project to decay by 50%. Fitted as an exponential decay curve over my actual weekly commit counts. My love has a measurable half-life, and it is shorter than a gym membership. The chart also includes: BPM — beats per month. One commit, one heartbeat. The 2 AM index — [26]% of my commits happen between midnight and 5 AM. That is not a schedule. That is love. Ward census — [3] alive, [6] flatlined, [1] critical. Longest flatline — [ crypto-tracker ], silent for [2.2 years], built at the exact top of the market. And then — the part I'm proudest of — the app doesn't let you just feel bad . Every flatlined patient has a red button: ⚡ DEFIBRILLATE Pressing it opens a revival pledge on Solana : a memo transaction, signed with your own wallet, containing a vow to ship at least one commit to that repo within 7 days. It's permanent, timestamped, and
Arya Koste
2026-07-13 11:21
👁 9
查看原文 →
Dev.to
Day 134 of Learning MERN Stack
Hello Dev Community! 👋 It is officially Day 134 of my software engineering marathon! Today, I successfully extended the layout grids of my MERN Stack capstone e-commerce application, Sprintix , by implementing fully responsive feature banners, newsletter hooks, and a clean global footer! ⚛️🛡️📬 A premium storefront relies heavily on trust anchors and consistent site-wide navigational structures. Today's focus was ensuring these terminal layers look flawless across all viewport breaking thresholds. 🛠️ Deconstructing the Day 134 Interface Terminal As captured in my local hosting environments within "Screenshot (301).jpg" and "Screenshot (302).jpg" , the system layout introduces high-fidelity structural blocks: 1. Trust Policy Infrastructure Positioned a 3-column micro-service layer layout framing crucial customer success policies (Easy Exchange, 7 Days Return, 24/7 Support). Balanced standard tracking font sizes and vector alignments to maintain optimal layout readability. 2. Immersive Newsletter Conversion Segment Engineered an engaging email onboarding banner using rich layered visual configurations. Integrated a responsive inline input element paired with an absolute action button to ensure the container shifts scales perfectly when transitioning down to mobile form factors. 3. Consolidated Multi-Grid Footer System Look at "Screenshot (302).jpg" ! Structured a highly scalable flex-wrapping matrix containing: Brand Identity Columns hosting contextual descriptive descriptions. Navigational Routing Indexes pointing clearly to operational views (Home, About Us, Privacy Policy). Direct Touchpoints aggregating structural contact details. Finished off the grid matrix with a clean full-width divider row holding structural copyright information. 💡 The Technical Win: Designing for Fluid Responsiveness First When building high-traffic online stores, mobile responsiveness isn't a secondary polish step—it has to be native. Writing components with flexible flexbox wrapping, relat
Ali Hamza
2026-07-13 11:18
👁 7
查看原文 →
Dev.to
Old projects
I recently found an old project I built with a friend around 2017–2018: a perk calculator for the game Firefall. The application allowed players to browse perks by category, drag them into a build, track the available perk points and automatically filter incompatible options based on the selected class. Looking at the code today, there are many things I would structure differently. The JavaScript could be better organised, responsibilities could be clearer, and the overall architecture would benefit from more modern practices. Still, I decided to preserve it as it is. Older projects are useful reminders that progress is not only visible in the technologies we use, but also in how we model problems, organise code and make technical decisions. It is not a showcase of how I would build the same application today. It is a snapshot of how I approached a real problem at that point in my career. Repository: https://github.com/lksvn/firefall-perk-calculator
LkSvn
2026-07-13 11:17
👁 7
查看原文 →