今日已更新 280 条资讯 | 累计 36295 条内容
关于我们

标签:#m

找到 10812 篇相关文章

AI 资讯

Presentation: Write-Ahead Intent Log: A Foundation for Efficient CDC at Scale

Vinay Chella and Akshat Goel discuss the challenges of running traditional CDC across heterogeneous databases during peak order traffic. They explain how Debezium hit limits under high load and share how they built Write-Ahead Intent Log (WAIL) - a custom architecture that utilizes a dumb producer proxy and a smart consumer pattern to cleanly separate the intent from the state payload. By Vinay Chella, Akshat Goel

2026-06-18 原文 →
AI 资讯

Amazon’s Kindle Colorsoft bundle is almost half off for Prime Day

Amazon’s Kindle Colorsoft Essentials Bundle is on sale for $182.97 (originally $334.97) as an early Prime Day deal, the lowest price we’ve seen for the combo. Unlike other Kindles, the Colorsoft’s color E Ink screen is great for comic books and graphics novels, illustrated books, or just perusing book covers while deciding what to read […]

2026-06-18 原文 →
AI 资讯

Struggling with PDF scanned nested tabels to html/md/json conversion

For the past few days, I've been trying to parse a PDF (scanned and text based) which has the same contents. PDF has nested tables Tables start at one page and end at another Currently I have been using (docling)[ https://docling-project.github.io ] to help me out with this text and convert it to the formats I require. I have a few limitations that I have a limit of 30s per page (2 minutes for the 4 paged pdf). And the biggest limitation is that I have to optimize it for the CPU. It has to run at a maximum of around 30s per page on the CPU . I have been trying a lot but docling is always failing at figuring out the table breaking in between two pages, and one single table and information that spills out from one page to another, is created into different tables by docling. How do I resolve this ? I am not being able to find a solution that works well within my given constraints better than docling currently. I've tried PyMuPDF, I've tried camelot as well. Camelot gave very nice results in converting to CSV, but it fails when nested tables come into the picture. I even tried to integrate camelot + docling into a hybrid pipeline but that also failed with my PDF with nested tables. Has anyone faced this problem before? Does anyone know of resources that could help me out with this problem? Any recommendations? Anything? :sob: submitted by /u/Kakarot_DB [link] [留言]

2026-06-18 原文 →
AI 资讯

Generative AI vs Agentic AI vs AI Agents [2026 Compared]

Originally published at kunalganglani.com — read it there for inline code, hero image, and live links. Generative AI vs agentic AI vs AI agents. Three terms, used interchangeably by people who should know better, burning engineering budgets across the industry in 2026. Generative AI refers to models that produce new content — text, images, code — from a prompt. AI agents are software systems that wrap those models with planning, memory, and tool use to pursue goals autonomously. Agentic AI is the broader paradigm: orchestrated systems of agents, workflows, and decision-making that operate with minimal human oversight. Getting these distinctions wrong doesn't just lose you a Twitter argument. It determines whether your production system costs $500/month or $50,000. Every quarter, someone on a leadership team says "we need to go agentic." What they usually mean is one of three completely different things. And the architecture you pick for each one has wildly different implications for cost, latency, reliability, and maintenance burden. I've watched teams burn entire quarters building autonomous agent systems when a well-tuned prompt engineering pipeline would have shipped in a week. That's not a hypothetical. I watched it happen twice in 2025. This post cuts through the buzzword soup. I'll define all three paradigms with concrete technical distinctions, show you how they map to real production architectures, and give you a decision framework for picking the right one. What Is Generative AI? The Engine, Not the Vehicle Generative AI is the foundation layer. It's a large language model (or image model, or audio model) that takes an input and produces new output. GPT-4, Claude, Gemini, Llama — these are all generative AI. You send a prompt, you get a completion. That's it. The critical thing to understand: generative AI is stateless by default . Each API call is independent. The model doesn't remember what you asked five minutes ago. It doesn't plan a sequence of steps.

2026-06-18 原文 →
AI 资讯

Fencing a node that doesn't know it's dead: pgrac build log #2

pgrac is an open attempt to rebuild Oracle RAC's core machinery (shared-everything storage, multiple active nodes all writing one database, a cluster-wide change number) on top of PostgreSQL 16. Build log #1 laid out the four problems that fight back. This one is about the problem that turns a node failure into silent data corruption, and the first, deliberately modest, layer pgrac ships against it. The failure mode In a shared-nothing cluster an evicted node is mostly harmless: it owns its own disks, so the cluster routes around it. In a shared-everything cluster the same event is dangerous, because every node writes the same storage. Picture the classic split: node 2 misses heartbeats, the cluster declares it dead and remasters its work elsewhere, but node 2 is not actually dead. It is frozen on a long GC pause, or its interconnect NIC flaked, and it is about to wake up and finish the write it started. Now two nodes believe they own the same blocks, and shared storage will accept both writes. That is not a crash. It is corruption you find three days later. Oracle RAC's answer is I/O fencing: before remastering a dead node's resources, you make certain it can no longer touch the storage, with STONITH, SCSI-3 persistent reservations, or a hardware watchdog. The node is fenced at a layer below its own software, because the whole point is that you cannot trust the dead node's software to behave. That hardware layer is real work, and it is not what pgrac built first. What it built first is the layer above it: an in-process cooperative write-fence, now default-ON. The rest of this is precise about what that does and does not buy you, because "we have fencing" is the kind of claim that is worth less than nothing if it is overstated. A fence needs an authority everyone can agree on You cannot fence on local opinion, because the whole problem is that the dead node disagrees about being dead. Authority has to live on durable, shared, quorum-backed storage. pgrac writes a sm

2026-06-18 原文 →
AI 资讯

I Went Looking for the Basis of 'N Characters Per Minute Is Fast' — There Wasn't One. Setting Read-Aloud Thresholds Honestly

📝 Originally published in Japanese on Zenn. This is the English version. Canonical: https://zenn.dev/uya0526_design/articles/satellite3_metrics-rationale 📚 This is satellite article #3 in my "Read-Aloud Speed Meter dev log" series. For the whole picture, see the main article . Where This Sits The read-aloud speed meter converts speaking speed into an evaluation label like "slightly fast," and stagnation rate into one like "few." Those labels ultimately become the foundation for Claude Haiku's feedback. So — on what basis did I draw the thresholds (the dividing lines)? This article digs into that "basis." The short answer from my research: I couldn't find a paper that defines an academic threshold for "N characters/min = fast/slow." This is a record of how I drew the lines honestly once I'd learned there was no firm basis. More than the metric numbers themselves, I believe being transparent about why I chose those numbers is what makes an evaluation app trustworthy. 💡 I'm an ex-Java engineer learning TypeScript in public. This one is mostly about design decisions. Why Obsess Over the "Basis"? An evaluation app passes judgment on the user: "your reading is slightly fast." Once you're passing judgment, if you can't explain "why we can say that," it's just guesswork. This app in particular passes the labels straight to Claude Haiku to generate coaching. If the foundational label has an unclear basis, the feedback built on top of it is a castle on sand. So I decided to nail down the basis for the thresholds first. Two things to research: The judgment basis for speaking speed (characters/min) The judgment basis for stagnation rate (the proportion of silence) As it turned out, these two had completely different kinds of basis. Speed Thresholds: No Academic Threshold → Draw From General Rules of Thumb What I found For speaking speed, I first looked for academic backing. Here's what I found: Speaking speed has traditionally been measured against mora count, but prior researc

2026-06-18 原文 →
AI 资讯

How I Have Build Memory That Actually Works for AI Coding

Most AI coding assistants do not really remember your project . They remember just enough to be dangerous . They see the latest prompt, skim a few files, improvise, and then forget the reasoning that made the answer useful five minutes ago. That is fine for toy demos. It breaks down fast inside a real software codebase. In Knotic I take an harder line . Instead of treating memory like a chat log with extra lipstick, I treat memory as infrastructure . Project knowledge is separated from session knowledge. Source material is separated from condensed understanding . Old context is compressed instead of blindly dragged forward. The result is a system that feels less like autocomplete with a caffeine habit and more like an AI engineering partner that can stay oriented over time. If you care about AI coding assistant memory , context engineering , persistent project memory , or long-term memory for software development , this is the part worth paying attention to. The Real Problem With AI Memory in Coding Tools The average AI IDE has the same failure mode . It looks smart on the first turn and shaky on the fifth . Why? Because software work is not just about answering the latest question. It is about carrying forward constraints, architecture, naming conventions, decisions, tradeoffs, dead ends, file relationships , and the exact context of the change in progress. When an assistant does not separate those layers, everything gets mixed together . Stable project facts sit next to temporary tool output. Important decisions compete with random noise. The model burns tokens re-reading the same files, or worse, works from partial memory and starts making up the missing pieces . Knotic solves this by splitting memory into distinct layers , each with a clear job. That design choice sounds simple. In practice, it changes everything . Knotic Does Not Use One Memory. It Uses Three. Knotic's memory model is built around three different kinds of context . The first is long-term projec

2026-06-18 原文 →
AI 资讯

I spend more time gathering context than completing coding tasks

I've been an engineer for almost 9 years, and I know from experience how much coding has changed over the years. Right now Im working in a big blockchain company and honestly I feel pretty exhausted. BUT NOT FROM THE TASKS I EXECUTE. I think with AI now, my work is more like being a human API. Lol. I got to slack, emails, JIRA and zoom calls to interact with people and gather all the context needed in order to make sure that when I will use AI the results will be relevant and accurate. And I feel that this is actually draining me. And i realized that this because every time we open PRs and it is about time to review things, CIs are freaking failing everywhere and then I have to go back and forth with people on slack to get the missing context. And all that even if we have already done scoping, architectural decisions. I feel we rush so much to deliver things fast, due to the AI-speed pressure, that is causing all this. I actually found many articles online talking about this. Anthropic also did their own index for checking if the fatigue is real from AI usage. I linked a medium article that resonated with me on the topic. Are you also facing this issue at your job? If so, how are you dealing with this, apart from taking more walks at the park lol. submitted by /u/LeopardAfter493 [link] [留言]

2026-06-18 原文 →