今日已更新 228 条资讯 | 累计 30272 条内容
关于我们

AI 资讯

AI人工智能最新资讯、模型发布、研究进展

15996
篇文章

共 15996 篇 · 第 700/800 页

Dev.to

From Pills to Pixels: Building an Intelligent Home Pharmacy Manager with YOLOv8 and CLIP 💊✨

We’ve all been there: staring at a messy medicine cabinet, wondering which box is for allergies and which one expired in 2022. In the world of Computer Vision and AI Healthcare , digitizing physical assets is a classic challenge. Today, we're building a "Medicine Box Expert"—a sophisticated pipeline that uses YOLOv8 for precision detection and OpenAI CLIP for multimodal understanding to turn a pile of pills into a searchable digital database. By the end of this tutorial, you'll understand how to bridge the gap between raw pixels and structured medical data. We are moving beyond simple classification; we are building a robust system capable of handling complex lighting, varied angles, and the tiny typography common in pharmaceutical packaging. The Architecture: A Multi-Stage Vision Pipeline To achieve high accuracy, we don't rely on a single model. Instead, we use a "Detect-Extract-Embed" workflow. graph TD A[User Uploads Image] --> B[YOLOv8: Box Detection] B --> C{Box Found?} C -- Yes --> D[Crop & Preprocess] C -- No --> E[Error: No Box Detected] D --> F[Tesseract OCR: Text Extraction] D --> G[OpenAI CLIP: Visual Embedding] F & G --> H[SQLite Query: Semantic Search] H --> I[Result: Drug Info & Dosage] Prerequisites Before we dive into the code, ensure you have the following tech_stack installed: YOLOv8 : For real-time object detection. OpenAI CLIP : To handle semantic image-text matching. Tesseract OCR : For reading the fine print on the boxes. SQLite : To store and query our medicine metadata. pip install ultralytics transformers torch pytesseract Step 1: Detecting the Medicine Box with YOLOv8 First, we need to locate the medicine box within the frame. A generic YOLOv8 model (like yolov8n.pt ) is surprisingly good at detecting "books" or "cell phones," but for the best results, you should fine-tune it on the Open Images Dataset specifically for "Box" or "Medical Packaging." from ultralytics import YOLO import cv2 # Load the model model = YOLO ( ' yolov8n.pt ' ) def

wellallyTech 2026-06-03 08:40 👁 7 查看原文 →
Dev.to

CIFSwitch - CVE-2026-46243

Just released an open-source bash checker for CIFSwitch (CVE-2026-46243) — the 19-year-old Linux kernel LPE disclosed last week that lets any unprivileged local user get root by abusing the CIFS/SPNEGO upcall path. The script runs on bare-metal, VMs, and inside containers, and is CI/CD-friendly with JSON output and clean exit codes. It checks: ✅ Kernel version against patched thresholds (6.18.22 / 6.19.12 / 7.0+) ✅ cifs-utils presence and exploitable version ✅ CIFS kernel module load state and blacklist status ✅ Unprivileged user namespace sysctl (the pivot point for the exploit) ✅ Active request-key cifs.spnego rules ✅ SELinux / AppArmor enforcement ✅ Container capabilities (CAP_SYS_ADMIN) ✅ Kernel symbol verification for the fix commit Outputs human-readable or JSON for SIEM ingestion. Exit 0 = safe, exit 1 = action needed — drop it straight into a pipeline. CIFSwitch is the fourth Linux LPE in under six weeks (after Copy Fail, Dirty Frag, and Fragnesia). If you're running multi-tenant Linux, CI runners, or container build farms, now is a good time to audit. I have also updated the cve_checks.conf in my my K8s-container_escape_audit toolkit to detect this issue.

Liam Romanis 2026-06-03 08:33 👁 11 查看原文 →
Dev.to

Microsoft MAI-Thinking-1 & MAI-Code-1-Flash: Developer Guide to 7 New MAI Models

Microsoft launched seven new in-house AI models at Build 2026 on June 2, 2026, marking the company's most significant push yet to build its own frontier AI stack independent of OpenAI. The centerpiece is MAI-Thinking-1, Microsoft's first large-scale reasoning model, built from scratch on clean commercially licensed data using a sparse Mixture of Experts architecture. Alongside it: MAI-Code-1-Flash, a 5-billion-parameter coding model that outperforms Claude Haiku 4.5 by 16 percentage points on SWE-Bench Pro while using 60% fewer tokens on complex tasks. This is the complete developer guide to all seven MAI models, their specs, benchmarks, deployment paths, and what they mean for the AI development ecosystem. Why Seven Models at Once? The strategic context matters. For three years, Microsoft's AI product surface — GitHub Copilot, Azure AI, Bing Chat, Microsoft 365 Copilot — ran almost entirely on OpenAI models. The Build 2026 announcement is Microsoft's public declaration that it is building a parallel, proprietary model stack. Every new MAI model is trained from scratch using "clean and appropriately licensed data, without distillation from third-party models" — language that directly addresses the intellectual property concerns that have accompanied third-party model licensing. The distribution strategy is equally deliberate. Microsoft is not routing MAI models exclusively through Azure. MAI-Thinking-1 and MAI-Code-1-Flash are available via Fireworks AI, Baseten, and OpenRouter — three infrastructure providers that collectively reach developers who explicitly do not want cloud vendor lock-in. This signals a platform-first posture: Microsoft wants MAI to become a model ecosystem, not just an Azure feature. MAI-Thinking-1: The Reasoning Flagship MAI-Thinking-1 is Microsoft's answer to Claude Opus 4.x and GPT-5.5 on the reasoning side of the model spectrum. The architecture is a 35-billion-parameter active / approximately 1-trillion-parameter total sparse Mixture of Ex

Anup Karanjkar 2026-06-03 08:23 👁 9 查看原文 →
Dev.to

Claude Opus 4.8 shipped today. Here is what the launch post does not say about why your agents will feel different tomorrow.

Claude Opus 4.8 shipped today. The benchmarks are a distraction — here is what actually changes about how your agents run tomorrow. Anthropic announced Claude Opus 4.8 at 16:00 UTC on June 3, 2026. The launch post leads with the usual benchmark deltas: SWE-bench Verified up 4.1 points, GPQA Diamond up 2.9, TAU-bench tool-use up 6.4. There is a chart. There is a marketing line about "the most capable agentic model we have ever shipped." If you stop reading there, you will miss the three things that will change how your production agents behave starting tomorrow. I have spent the morning re-running our internal agent harness against Opus 4.8 and reading the model card line by line. Two of the three changes are improvements. One of them is a silent regression that will bite anyone who pinned the model ID. Here is the full picture. What 4.8 actually changes The model card and release notes ship three changes that the launch blog post does not foreground: Cache-aware routing inside long agentic loops. The 4.7 router treated every tool-call cycle as a fresh planning step. 4.8 keeps an internal trace of which cache breakpoints were hit on the previous step and biases the next plan toward extending those traces. In agent harnesses that already use prompt caching aggressively (Claude Code, the Agent SDK with cacheControl: "ephemeral" on the system prompt), cache hit rates jumped from a measured ~46% on 4.7 to ~71% on 4.8 across a 30-step coding loop. The 200k context window now actually behaves at 200k. Anthropic published a needle-in-a-haystack chart in the model card going out to 200,000 tokens. The 4.7 chart got noticeably worse past ~140k tokens; the 4.8 chart is flat. This sounds like a benchmark thing. It is not. It changes the cost equation for "just stuff everything in context" patterns that 4.7 quietly punished by degrading accuracy. claude-opus-4-7 was not aliased. The launch shipped a new model ID — claude-opus-4-8 — and the previous ID is still callable. But if y

LayerZero 2026-06-03 08:11 👁 6 查看原文 →
Dev.to

I Wrote 40 Lines of Python to Beat Tokyo Salaries from Rural Japan: Furusato Nozei + Utility Defense for Remote Side-Hustlers (2

⚠️ この記事はアフィリエイト広告(プロモーション)を含みます。リンク先で発生した収益の一部が運営者に支払われますが、読者の購入価格には一切影響ありません。 If you work remote from rural Japan, by the end of this article you'll have two runnable Python scripts: one that computes your exact furusato-nozei (hometown tax) ceiling from your real side-income, and one that scores your electricity contract against your actual kWh log so you stop overpaying. No spreadsheets, no "consult a tax accountant" hand-waving. Copy, run, save money tonight. Result from my own 2025 numbers: ¥41,000 of furusato-nozei reward goods for a net cost of ¥2,000, plus ¥28,400/year shaved off my power bill after switching plans. Total ≈ ¥67,400 recovered, and because I work from home in Niigata, my commute cost to earn it was literally ¥0. The trap: side income breaks the "simple" furusato nozei calculator Every portal (Satofuru, Rakuten Furusato, Furunavi) shows a slider that estimates your ceiling from salary alone. The moment you add freelance/blog/ Kindle income, that slider lies to you. In 2024 I trusted it, donated ¥52,000, and ¥9,000 of it fell outside the deductible ceiling because my side income pushed me into a different residual-tax bracket. That ¥9,000 was just a donation — no tax back. The real ceiling depends on your total taxable income (salary + side hustle minus expenses) and the resident-tax (juminzei) cap of roughly 20% of your income-based resident tax. Here's a calculator that actually folds in side income. It uses Japan's 2026 progressive income-tax brackets. # furusato_ceiling.py — Python 3.9+ from dataclasses import dataclass # 2026 national income tax brackets: (upper_bound_yen, rate, deduction_yen) BRACKETS = [ ( 1_950_000 , 0.05 , 0 ), ( 3_300_000 , 0.10 , 97_500 ), ( 6_950_000 , 0.20 , 427_500 ), ( 9_000_000 , 0.23 , 636_000 ), ( 18_000_000 , 0.33 , 1_536_000 ), ( 40_000_000 , 0.40 , 2_796_000 ), ( float ( " inf " ), 0.45 , 4_796_000 ), ] @dataclass class Taxpayer : salary_income : int # after salary-income deduction (給与所得) side_profit : int #

スシロー 2026-06-03 08:08 👁 10 查看原文 →
Dev.to

[iOS] Why Passthrough MP4 Export Failed for iPhone Videos

A user picks a video from Photos. The app turns that video into a file the server can accept. Then it uploads the file. On the surface, this sounds like a simple upload flow. In practice, iOS makes you answer a much more specific question: How do you reliably turn a video from the user's Photo Library into an uploadable MP4 file? At first, I used AVAssetExportPresetPassthrough . It looked like the right default. It preserves the original media as much as possible, avoids unnecessary re-encoding, and is fast when it works. No quality loss, less CPU usage, less battery cost. But some videos failed during export. The error was usually in the AVFoundationErrorDomain Code=-11838 family. Apple describes this error as operationNotSupportedForAsset : an operation was attempted that is not supported for the asset. At first, I suspected an iCloud download issue, a Photos permission edge case, or something retryable inside PHImageManager . That was not the real problem. The actual problem was more fundamental: I was trying to write an iPhone MOV asset into an MP4 file using a passthrough export. The Original Flow The original code was roughly shaped like this: PHImageManager . default () . requestExportSession ( forVideo : asset , options : options , exportPreset : AVAssetExportPresetPassthrough ) { exportSession , info in exportSession ? . outputURL = outputURL exportSession ? . outputFileType = . mp4 exportSession ? . exportAsynchronously { // upload } } The intention was reasonable: Ask Photos for an export session for the PHAsset . Use AVAssetExportPresetPassthrough to preserve the original quality. Write the result as an .mp4 file for upload. Many videos exported successfully this way. That is what made the bug confusing. Some videos worked. Some did not. The hidden assumption was: Any iPhone video can become an MP4 file through a passthrough export. That assumption is not always true. iPhone Videos Are Usually MOV Files To a user, it is just "a video." Internally, an iPh

raykim 2026-06-03 08:07 👁 10 查看原文 →
Reddit r/artificial

If your AI agent can send emails, browse websites, or call tools, I want to test something with you

Most security tools for AI agents check one message at a time. Arc Gate tracks the whole conversation. That matters because the attacks that actually work in production don’t happen in one message. They happen across 8 turns. Each one looks clean. By the time the payload arrives your agent is already primed to execute it. I built Arc Gate using a geometric framework from my own research to detect adversarial behavioral drift across a full session — not just flag individual messages. When a conversation starts drifting toward something dangerous, it catches the pattern before the attack completes. I’m looking for 3 teams running real agents to test it against actual workflows and tell me where it breaks. Not chatbot wrappers. Agents with real tool access. Browser use, email actions, MCP servers, internal copilots, workflow automation. No charge. No sales call. Just feedback from people close to production. Comment or DM me if that’s you. Platform: https://bendexgeometry.com GitHub: https://github.com/9hannahnine-jpg/arc-gate Demo: https://web-production-6e47f.up.railway.app/demo submitted by /u/Turbulent-Tap6723 [link] [留言]

/u/Turbulent-Tap6723 2026-06-03 07:34 👁 5 查看原文 →
Reddit r/artificial

How does AI follow ethical guidelines in Data Collection?

Hot take: if I wanted to gather data via the internet, and I’m writing scripts/code to speed up the process, I have to follow some basic rules (ie look at the sitemap, find relevant robots.txt, follow that websites preference and rules). But it seems any AI-agent I’ve used does not give af about rules and limits, and is totally cool building me a scraper that will perform hundreds of thousands of requests without regards to the website owner’s preference. Given it’s widely known you can use AI for simple coding tasks I can easily see a future where ordinary individuals are operating their own scrapers. Especially in gathering high-value information that “seems easy to get” like google search rankings, or job data. This creates an obvious nightmare for Google, ATS platforms, and just about every website on the internet if everyone and their mother starts spinning up Playwright sessions in Python. I’m deadset on this being a responsbility of AI providers (anthropic, open ai, anysphere, etc). But how are these companies supposed to balance this without implementing guardrails that heavily limit their products? Maybe this has been solved and someone can feed my curiosity. submitted by /u/TacoTuesdayX [link] [留言]

/u/TacoTuesdayX 2026-06-03 06:53 👁 5 查看原文 →
The Verge AI

God of War Laufey is coming to the PS5

Sony ended its big State of Play showcase with a major reveal: the next God of War. The new title is called God of War Laufey, and is once again developed by Sony's Santa Monica Studio. Currently, the game doesn't have a date, but it's coming to the PS5 whenever it does launch. This time, […]

Andrew Webster 2026-06-03 06:16 👁 5 查看原文 →
Dev.to

Next.js 16 Caching for E-Commerce: Cache Components, use cache, revalidateTag, and Fresh Product Data

Caching in e-commerce is never just about speed. A fast storefront is useful only if it still shows the right price, the correct stock level, and the right experience for the current customer. That is why caching in a Next.js storefront can be deceptively hard. Some data should be shared broadly and kept warm for SEO and performance. Some data should be refreshed often. Some should never be shared between users at all. Next.js 16 gives teams a much clearer toolbox for solving this problem with Cache Components, use cache , tag-based invalidation, and explicit cache lifetime controls. Used properly, these features let you keep pages fast without drifting into stale commerce data. In this guide, I will walk through a practical way to think about caching in a modern storefront and show how to combine use cache , cacheLife , and revalidateTag for real e-commerce use cases. Why Caching Is Harder in E-Commerce Than in a Typical Content Site On a standard marketing site, most content changes infrequently. If a page is cached for a few minutes or even a few hours, the business impact is usually negligible. Commerce systems work differently. The same product page may contain: stable product descriptions and category copy semi-dynamic data such as price, availability, shipping estimates, or promotion labels private data such as cart state, recently viewed items, or customer-specific pricing Treating all of that data the same way leads to one of two bad outcomes: You cache too aggressively and serve stale prices or availability. You disable caching everywhere and lose the performance benefits that help SEO and conversion. The better approach is to split your data by volatility and audience. The Three Cache Boundaries That Matter Most For most commerce projects, the cleanest mental model is to divide data into three groups. 1. Stable catalog content This is the part of the page that usually changes only when content editors or merchandisers update the catalog. Examples: product

uninterrupted 2026-06-03 06:00 👁 13 查看原文 →
The Verge AI

Remedy’s Control sequel launches in September

Control Resonant, the upcoming sequel to Remedy Entertainment's Control, will be released on September 24th, 2026, according to a trailer that premiered during Tuesday's PlayStation State of Play show. The trailer also gave a preview of the game's story, which stars Dylan Faden, who was imprisoned for much of the first game. He will be […]

Jay Peters 2026-06-03 05:55 👁 13 查看原文 →
Reddit r/artificial

What's the best AI image generator for fine art?

For the AI artists here who like creating painterly / impasto style work, what generators are you guys using? Just curious since I'm much more into that theme as opposed to realism, but it feels like most generators cater to realism now. submitted by /u/geekedprompts [link] [留言]

/u/geekedprompts 2026-06-03 05:55 👁 8 查看原文 →
Dev.to

A Curated List of Articles About Modern Software Testing

Software testing is changing quickly. Teams are dealing with faster release cycles, more AI-assisted development, more complex browser behavior, and higher expectations around product quality. I collected a few practical articles that cover different parts of modern QA, test automation, developer workflows, and testing strategy. Recommended reads How to Test AI Agents for Tool Use, Memory, and Recovery Paths A practical framework for testing AI agents for tool use, memory retention, retries, and recovery paths, with concrete strategies for QA and engineering teams. How to Evaluate a Test Automation Tool for Shadow DOM, iframes, and Other Hard-to-Test UI Surfaces A practical buyer guide for evaluating test automation tools for shadow DOM testing, iframe testing, resilient selectors, and dynamic UI edge cases. How to Reproduce a Flaky Browser Test with Video, Logs, and Network Traces A practical workflow to reproduce a flaky browser test using video, logs, and network traces, then turn intermittent failures into repeatable bug reports. Endtest Review for Small QA Teams: Where Editable Test Flows Save the Most Time A practical Endtest review for small QA teams focused on editable test flows, maintainable test steps, and where no-code QA automation actually saves time. Editable Test Steps vs Generated Test Code: Which Holds Up Better After UI Changes? A practical comparison of editable test steps vs generated test code for UI change resilience, maintenance overhead, debugging, and team handoff, with guidance for QA and engineering leaders. Managed QA Services vs Staff Augmentation: What Changes in Ownership, Speed, and Cost A practical comparison of managed QA services vs staff augmentation, focusing on ownership, ramp time, communication overhead, cost, and maintenance risk. Automation Payback Period: How Long Does QA Test Automation Take to Break Even? Learn how to estimate the test automation payback period, model QA ROI, account for maintenance cost, and identify wh

Antoine Dubois 2026-06-03 05:49 👁 11 查看原文 →
Dev.to

"Todo mundo sabe mais que eu?" Como encarei a síndrome do impostor na faculdade de Computação

Em 2023, eu arrumei as malas e saí de Recife, minha terra natal, para morar no interior do Mato Grosso. O motivo? Cursar Ciência da Computação. Eu sabia que enfrentaria um choque cultural, o calor do Centro-Oeste e a saudade de casa. O que eu não imaginei é que o maior desafio seria a sensação constante de que eu era uma fraude no meio da minha própria sala de aula. Se você estuda ou trabalha com tecnologia, provavelmente já passou por isso. Você senta na primeira semana de aula e parece que metade da turma já programa desde os 12 anos, fala termos técnicos que parecem outra língua e discute sobre ferramentas que você nem sabia que existiam. Hoje, na reta final da graduação, focando meus estudos em Back-end e Dados, quebrando a cabeça com Go, Python e SQL, eu olho para trás e vejo o quanto essa cobrança silenciosa quase me paralisou. Se você está sentindo que entrou no curso errado ou que todo mundo corre a 100 km/h enquanto você ainda está engatinhando, pega um café e vem ler este papo reto sobre como lidar com a tal da Síndrome do Impostor. O "efeito vitrine" e a falsa ilusão de que somos os únicos perdidos Na faculdade de TI, a comparação é quase inevitável. A gente abre o LinkedIn e vê o colega conseguindo estágio internacional; abre o GitHub e vê códigos impecáveis; na sala de aula, sempre tem quem responda às perguntas do professor antes mesmo dele terminar de falar. O erro está em achar que o ritmo do outro deve ditar o seu. Na Computação, as pessoas vêm de bagagens completamente diferentes. Quem já sabia programar antes da faculdade pode ter facilidade em Lógica de Programação, mas talvez sofra tanto quanto você quando chegar a hora de aprender Álgebra Linear ou Teoria da Computação. Sentir-se perdida não significa incompetência; significa apenas que você está no processo de aprender algo complexo. E adivinha? Todo mundo ali está tentando descobrir como as coisas funcionam, mesmo quem finge que sabe tudo. O que me ajudou a virar a chave Eu não acordei um dia

Raissa Cavalcanti 2026-06-03 05:42 👁 14 查看原文 →
Dev.to

Escudo

Privacy-First Personal Finance for iOS Your finances. On your phone. Nowhere else. A privacy-first personal finance app that connects your banks, brokerage, and investment accounts into one unified dashboard — entirely on-device, no backend, no account required, no subscription. View on GitHub At a Glance 🏦 Multi-source 🔒 100% On-device 📊 Full picture Banks, Revolut, Trading 212 and more No server. No account. Your data stays in your Keychain. Net worth, spending, investments — all in one place Screenshots Log Insights Budget Transaction Entry Settings About Escudo Built out of two frustrations: every decent finance app costs a monthly subscription, and none of them support Trading 212. Escudo connects your banks, brokerage, and investment accounts and gives you a single view of your net worth, spending, and investments — without your data ever leaving your phone. Key Features Net worth dashboard — aggregated balance across all accounts and investment portfolio P&L Unified transaction log — every account in one feed, auto-categorised Spending insights — breakdown by category and trends over time Budget tracking — per-category budgets with visual progress dials Multi-currency — EUR, GBP, USD with stored exchange rates Recurring transactions — template-based recurring transaction engine Shortcuts support — deep linking via escudo:// URL scheme Integrations Source Method Trading 212 REST API — portfolio, orders, dividends Revolut Enable Banking OAuth 2.0 Bankinter PT Enable Banking OAuth 2.0 SIBS SIBS Open Banking (PT market) CSV import Revolut & Bankinter statements All credentials live in the iOS Keychain — never in UserDefaults, never in iCloud, never on a server. Known Limitations Enable Banking does not expose credit card accounts — only bank accounts and transactions are available through the PSD2 API; credit card balances and transactions are not accessible No token auto-refresh for Enable Banking — manual re-auth when tokens expire Categorisation rules are hard

Henrique Peralta 2026-06-03 05:39 👁 10 查看原文 →