AI 资讯
Building Picturesque AI: one studio, 50+ models, and the plumbing nobody wants to maintain
One creative studio for images, video, music, audio, editing, upscaling, and motion control. 50+ models, one credit balance. This is mostly about how we built it and what went wrong along the way. The problem (from a dev perspective) The models are good now. That's not really the issue anymore. The issue is everything around them. Different providers, different UIs, different billing. No shared history across modalities. No easy way to go from "generate image" to "animate it" to "add music" to "upscale" without opening four tabs. We wanted one place where you could actually finish something. What the product is Picturesque has a few main pieces: Studio - tabs for image, video, audio, edit, motion control Projects + Explore - save your work, browse what other people made Workflows - node canvas where you chain models together and run the pipeline in one go Director - an agent that plans multi-step creative work, quotes credits, and runs generations for you The studio covers a lot on its own. 4K images, cinematic video with audio, Suno music, ElevenLabs TTS, Topaz upscaling, motion control, talking avatars. The annoying engineering showed up once we tried to make all of that feel like one product instead of a folder of integrations. Stack (kept boring on purpose) Frontend is React, TypeScript, Vite, React Router. Backend is Node + Express. Socket.IO for real-time updates. Supabase for Postgres and auth. S3-compatible storage for outputs and uploads. For the actual model calls we built a service layer that normalizes inputs, maps our internal model IDs to provider APIs, and handles retries/errors in one place. Media stuff runs through FFmpeg and Sharp. Nothing fancy. When you're wiring up dozens of models with different schemas and pricing rules, you don't want your infra adding more chaos. We also refactored the backend out of a single 7,700-line server.js into routes + services. Painful refactor. Would do it again immediately. The unglamorous part: 50 models, one UI
AI 资讯
How to Create a Skill in Claude Code
This is a cross-post — the original (and any updates) live at broke2builtai.com . The first time I watched Claude Code reach for a skill I hadn't told it to use — read a folder, run the script inside it, and hand back the finished thing — the difference from a slash command finally landed. A slash command waits for you to type it. A skill waits for the situation . Claude decides. That one shift is the whole feature, and building one takes about five minutes once you know where the file goes. Here's the entire thing end to end, including the one gotcha that decides whether your skill ever actually fires. What a Skill actually is A Skill is a folder with a SKILL.md file inside it. The Markdown holds instructions; the YAML frontmatter at the top holds a name and a description . That description is doing the most important job in the whole file: Claude reads it to decide, on its own, whether the current task warrants invoking the skill. Nothing else you write matters if the description doesn't get you picked. That's the mental model to hold onto: a custom slash command is a prompt you trigger by typing /name ; a skill is a procedure Claude triggers when the context matches. Same reusable-instructions idea, opposite trigger. Where the file goes Two locations register, exactly like commands and subagents : Project skill — .claude/skills/<skill-name>/SKILL.md inside the repo. Committed, so your whole team gets it. Personal skill — ~/.claude/skills/<skill-name>/SKILL.md in your home directory. Follows you across every project on your machine. Each skill is its own folder, and the folder name should match the name in the frontmatter. A loose SKILL.md sitting somewhere else won't be picked up. The minimum viable skill Create the folder and the file: .claude/skills/pytest-runner/SKILL.md Then write the two-part file — frontmatter, then body: --- name : pytest-runner description : " Run, generate, or debug pytest tests for this project. Use when the user asks to run the test su
AI 资讯
Our Journey to GSSoC 2026: Omnikon's Repository Has Been Selected! 🎉
Open source has always been at the heart of what we do at Omnikon. Today, we're excited to share a milestone that means a lot to our entire community. Our repository, maintained by Sourabh, has been officially selected for GirlScript Summer of Code (GSSoC) 2026. For us, this isn't just another achievement—it's a step toward building a stronger open-source ecosystem where students and developers can learn, collaborate, and create meaningful software together. About Omnikon Omnikon is a student-led open-source organization focused on building high-quality developer tools, educational resources, and community-driven projects. Our mission is simple: Build impactful open-source software. Help new contributors get started. Create projects that solve real problems. Foster a welcoming developer community. Every repository we build is designed with collaboration in mind, making it easier for contributors of all experience levels to participate. What GSSoC Means GirlScript Summer of Code is one of India's largest open-source programs. Every year, thousands of contributors participate by solving issues, improving documentation, fixing bugs, and implementing new features across selected repositories. Being selected means our project will become part of this collaborative ecosystem, giving contributors an opportunity to make meaningful contributions while learning industry-standard development workflows. A Special Thanks This achievement wouldn't have been possible without Sourabh, who maintained and prepared the repository throughout the selection process. A huge thank you to everyone who contributed ideas, reviewed code, reported issues, improved documentation, and supported the project. Open source is never the work of one person—it grows because of a community. What's Next? We're preparing the repository for contributors by: Organizing beginner-friendly issues. Improving documentation. Creating contribution guides. Enhancing project structure. Mentoring new contributors thro
AI 资讯
Chrome Built-In AI APIs: A Hands-On Guide to Language Detection, Translation, Summarization and Writing Assistance
Introduction Chrome's Built-In AI APIs allow applications to perform selected AI workloads directly within the browser. Unlike traditional AI integrations, developers do not need to deploy or operate model infrastructure. This guide walks through the major APIs currently available. Getting Started: API Availability and Chrome Flags Chrome's Built-In AI APIs are at different stages of maturity. Some APIs are available in stable Chrome, while others remain experimental. The required setup therefore depends on the API you want to test. Available in Chrome Stable The following APIs are available in stable Chrome on supported desktop devices: Language Detector API Translator API Summarizer API These APIs do not require experimental flags for normal use in supported Chrome versions. The Prompt API has different availability requirements depending on whether it is used from a web page or a Chrome Extension. Check the current Chrome documentation for the environment you are targeting. Experimental APIs The Writer, Rewriter, and Proofreader APIs remain experimental and may require developer trials, origin trials, or Chrome flags for local development. Because these APIs are evolving, refer to the official Chrome documentation for the current setup requirements rather than relying on a static list of flags. Engineering recommendation: Use feature detection and availability() checks at runtime rather than relying on Chrome version numbers or assuming that a particular flag is enabled. Language Detector API Use cases: Dynamic localization Query routing Analytics Content classification Example const detector = await LanguageDetector . create (); const result = await detector . detect ( " Bonjour tout le monde " ); console . log ( result ); Architecture Notes Low latency Task-specific model Suitable for client-side execution Complete runnable example: Language Detector API on GitHub Gist Translator API Use cases: Localization Offline translation International applications Example
创业投融资
Guy who took photo of Jupiter with a Game Boy Camera and giant telescope publishes DIY tutorial
The latest adventure for the kooky camera looks skyward.
AI 资讯
The smartest model lost — and it just redrew the 2026 AI race
The most interesting model comparison of 2026 isn't a benchmark table. It's a product exec quietly changing the question everyone asks about models — and getting a completely different ranking as a result. Claire Vo (founder of ChatPRD, host of the How I AI podcast) ran a head-to-head between OpenAI's new GPT-5.6 lineup (Soul / Terra / Luna) and Anthropic's Claude Fable and Sonnet. The result was an upset: the most theoretically intelligent model, Claude Fable, lost to the one she could actually collaborate with, GPT-5.6 Soul. Here's what that upset actually reveals. She killed "vibes" — then bet 70% back on her own taste Tired of vibe-checking, Vo built a real benchmark across the work she does every day: writing PRDs, prototyping apps, debugging multi-step code, and talking to an agent. Scoring had two layers — an LLM-as-judge (she picked the harshest judge, GPT-5.5) and her own hand-graded "taste test," where she clicked through every artifact and wrote notes. Then the key move: she weighted the final score 70% her taste / 30% the machine. "It's my show. I trust my own taste more." That's the first insight. Benchmarks are getting more rigorous, but the final call is still human taste. The point of blind testing isn't to replace taste — it's to force it to be honest . Cover the labels, react to the work itself, then put your judgment back at the center. Theoretically brilliant vs. practically effective On raw intelligence, Fable is elite. But Vo's verdict is the sharpest line on models I've seen this year: Fable is theoretically hyper-intelligent. Soul is practically effective. She describes Fable as "an engineer who has never met a human." Precise to the point of pedantry — it scores every risk, hardens every edge. In one case it hardened a tool-calling loop so tightly that only one specific model could run it at all. It optimized itself into a corner. Soul's edge was the opposite: it gets out of its own head. Same stuck problem — she moved it to Codex, said "sto
AI 资讯
The Paintbrush Paradox: Why the Monolithic Era of AI Is Crumbling
Over the past week, two narratives have been colliding everywhere I look. On one side, there's panic. AI is expected to replace marketers, engineers, and entire categories of knowledge work almost overnight. On the other, there are quieter but far more consequential signals: enterprise teams discovering their AI infrastructure is burning through API budgets far faster than expected. This isn't because the underlying models are weak, but because the systems built around them are fundamentally inefficient by design. These aren't separate stories. They're the same failure showing up in different places. A conversation with another developer made that gap visible in real time. He argued that auditing a 150,000-line codebase requires feeding the entire repository into a model in one single, massive pass. It's still a common assumption in mainstream tech: that an LLM works like a giant biological brain that you must fully load with raw text before it can begin to think. But that assumption is already outdated. Modern AI systems don't scale through brute-force context. They scale through structure. And that shift changes everything. Key takeaways Bigger context windows did not solve AI. Treating a frontier model as a monolithic processor that re-reads an entire system on every query is wasteful, dilutes attention, and hides bugs under raw volume. ARC-AGI-3 makes the gap stark: frontier models scored under 1% on interactive reasoning tasks that untrained humans solve at nearly 100%. The gap is architecture, not memory. The teams pulling ahead treat the model as one narrow component inside a larger system: intelligent routing, task decomposition, retrieval, and only the minimum necessary context. The next advantage is not the biggest model or the longest prompt. It is the system designed around the model. Prompting was the first generation; systems architecture is the next. The Myth of the Infinite Context Window When context windows expanded into the hundreds of thousands o
AI 资讯
Salesforce Education Cloud: A Modern Alternative to EDA
Executive Summary The Salesforce Education Data Architecture (EDA) has served educational institutions well for over a decade as a free, community-supported managed package. However, with the 2023 launch of the reimagined Education Cloud—built natively on the Salesforce core platform—institutions now face a strategic choice about their CRM foundation . While EDA remains supported and continues to function effectively, Education Cloud represents a fundamental architectural shift that offers significant advantages in simplicity, scalability, and access to innovation . This paper examines why Education Cloud is demonstrably easier to implement and maintain compared to its predecessor, addressing the key differences in architecture, data model, and ongoing operations. 1. The Architectural Advantage: Built-In vs. Bolted-On 1.1 EDA: A Managed Package on Top of Salesforce EDA is a managed package installed on top of the Salesforce core platform . As a managed package, it creates additional layers of complexity: Installation and Updates: EDA requires separate package installations and updates that can lag behind Salesforce's native release cycle Namespace Conflicts: The managed package introduces its own namespace, potentially creating compatibility issues with other tools Translation Limitations: EDA's localization has documented issues, including a known problem where the Preferred Phone functionality fails when users switch to languages other than English Record Type Validation Bugs: Deactivating an account record type can block contact creation—a validation error that requires manual workarounds 1.2 Education Cloud: Native to the Core Platform Education Cloud represents a fundamentally different approach. Rather than being a package installed on Salesforce, Education Cloud is built directly on the Salesforce core platform . Key Advantages: No Package to Install: Education Cloud runs natively on the Salesforce core platform, eliminating the need for separate managed pack
科技前沿
Patch for Windows Defender 0-day could allow attackers to fill hard disk
The feud between NightmareEclipse and Microsoft shows no signs of resolving soon.
AI 资讯
The floatable, powerful Soundcore Boom 2 speaker is over half off
Bluetooth speakers with big sound and great features are hard to find for under $100, with most offerings being some variation of the same basic (and often small) design. Thankfully, through July 10th Woot has the Anker Soundcore Boom 2 on sale for $69.99, which is $20 cheaper than the discounted price at other retailers […]
产品设计
Allstate accuses Broadcom of auditing it because it quit VMware, CA
Broadcom accuses Allstate of dodging VMware audits.
开发者
Deploy a Dockerfile on Vercel
Yes, you heard it right, you can now run a Dockerfile on Vercel. Vercel was the go-to place where...
AI 资讯
Anthropic found a hidden space where Claude puzzles over concepts
The AI firm Anthropic has developed a technique that has given it the clearest glimpse yet at what’s really going on inside large language models as they answer questions or carry out tasks. What they found ranges from the mundane to the unnerving. Researchers at the company built a tool called the Jacobian lens (or…
科技前沿
Humanoid robots controlled by surgeons did world-first operation on live pigs
Preclinical trial is testing the feasibility of humanoid robots in surgery.
AI 资讯
Google will now tell you if an ad was made with AI
You can see if ads on Google Search, Google Discover, and YouTube were made or edited using AI from a new section in Google's "My Ad Center," as reported earlier by TechCrunch. The update, announced on Thursday, adds a "created or edited with AI" label under the "how this ad was made" tab. Users can […]
AI 资讯
Meta enters the crowded AI coding battle with Muse Spark 1.1
Meta's pitch to users is Spark's ability to handle large agentic workloads, fix bugs, and help with large code migrations — the kind of automation that enterprises are increasingly turning to AI companies to provide.
产品设计
Charles Hudson shares the common mistakes he’s seen after investing in 500+ startups
In this week’s episode of Build Mode, Isabelle Johannessen talks with Precursor Ventures' Charles Hudson about the headwinds facing early-stage founders today and the most common mistakes founders should avoid in order to get funded.
科技前沿
Judge doesn't like Elon Musk settlement with SEC, but says court can't block it
Judge reluctantly approves $1.5M settlement with SEC over Twitter stock violation.
AI 资讯
New York Times says OpenAI hid evidence in ChatGPT copyright trial
News publishers say OpenAI hid tools and datasets that could identify copyrighted journalism in ChatGPT outputs, escalating their lawsuit with a new motion for sanctions.
创业投融资
Slate Auto teams up with Crayola to color its EV truck
Slate has an answer for owners who have always want to drive a truck with bright crayon colors.