今日已更新 446 条资讯 | 累计 41220 条内容
关于我们

标签:#testing

找到 410 篇相关文章

AI 资讯

In a Few Weeks, I Won't Know How My Own App Works

I’ve been doing a lot of code reviews recently, and I’ve noticed something about my experience during the reviews. And what it means for code in general, but specifically for AI generated code. But before I do that, I want to take you back to the ’00s. A simpler time. Back then I started doing something I only read about in books – pair programming. Ok, that’s a lie. Everyone who grabbed someone and brought them to look at your code because it’s doing something weird – you’ve done pair programming. But this was pair programming for like 80% of my work. And of course, not just me. This links directly to my topic – our understanding of the code. And in the age before AI, where teams really knew their code, that would be the biggest grade I’d give for code understanding. Meaning, if you and I work both at the same time on the same task, and finish it. I’d say we’re both at the highest level of understanding of that feature – what it needs to do, how it works, what it depends on, what depends on it, and how it’s written and tested. From here that level of understanding is going to drop. If you’re part of a team, but didn’t work on that feature, you probably know it exists. If we have dependencies, you probably know about them. But the rest is either non-existent or close to it. If you’re on another team, that level of understanding drops, and if you’re on another project – you may not even be aware. That makes sense, but we’re not talking about that today, I want to focus on my capability to do a proper review. And I go back to understanding for that. In order for me to give you proper feedback, catch your mistakes and offer alternatives, I need proper understanding. The more understanding I have, I can review better and make better changes. So far, so good. It makes sense. Two more things make sense. First the use of tools for review. Every tool, from syntax analyzer to the smartest security checker, work according to patterns, not understanding. All the tools check fo

2026-09-09 原文 →
AI 资讯

Testability Is a Feature. Does Your Code Agent Know About It?

I always say that testability is a feature . It needs a customer – usually testers. We need to define what it means. And we need to build it in. The question is what happens to testability when someone else builds the code. Yup. Genie talk again. Let’s look at just 3 aspects of generated code, or any code really, that impact testability. But when our genie doesn’t get directions, it will skip those and cause us a big headache. 1. Complexity Of course, complexity affects testing. The more complex the code, the more work it takes to verify. More cases, more time – time that we don’t really have. My classic example is recursion. Recursion seems simple (in code lines), but it can hide all kinds of bugs in there. Edge cases galore. But that’s usually a function. That can be confined and handled. Today we’re generating systems. And they are as complex as the genie wants. Or the systems it was trained on. And those are not a model for simplicity. Complexity is the nemesis of testability. And unless we ask for simplicity, and make sure we got it, we pay for that in more expensive testing. 2. Observability Observability is a key part of testability. Without it, we may be able to operate the system, but may not see the impact of those operations. For example, if you had one POST API in the world that saves data in the database, and you want to check if it works, you’ll need to choose between looking inside the database (which may not be possible), or rely on the status code (which doesn’t tell you anything about the actual impact). Adding a GET to read helps, and presto – you’ve got observability. Of course, you need to ask for that GET API. In CRUD systems, you’d get that API, I’m not worried. Even in generated code. The problem starts with the not-so-intuitive stuff. What gets logged and where. Knowing how long it takes to see the result, and maybe how much time to wait between operations and states. The more the system exposes state and data, the easier it is to understand

2026-09-09 原文 →
AI 资讯

Batch Transaction - Testcases

Automated test coverage for the Batch transaction feature (XLS-56), grouped by the invariant or execution mode each test exercises with 189 tests. Category Test Count Execution Modes 42 Multi-Account & Multi-Sign 23 Tickets, Replay & Metadata 7 Vault, Loan & Transaction Types 24 Signature & Structural Validation 10 Security & Adversarial 56 Cross-Feature Interactions 27 Total 189 1. Execution Modes AllOrNothing Test Batch allornothing all payments succeed Test Batch allornothing submit batch multiple times Test Batch allornothing one payment fails Test Batch allornothing all payments fail Test Batch allornothing mixed transaction types Test Batch allornothing fee calculation Test Batch allornothing max inner transactions Test Batch allornothing more than max inner transactions Test Batch allornothing cash same check multiple times Test Batch allornothing fail and then succeed Test Batch allornothing with tickets Test Batch allornothing deep rollback on late inner failure Test Batch allornothing value conservation OnlyOne Test Batch onlyone first succeeds Test Batch onlyone first fails second succeeds Test Batch onlyone all fail Test Batch onlyone offer priority Test Batch onlyone max inner transactions Test Batch onlyone more than max inner transactions Test Batch onlyone cash same check multiple times Test Batch onlyone fail and then succeed UntilFailure Test Batch untilfailure all succeed Test Batch untilfailure stops at first Test Batch untilfailure stops at second Test Batch untilfailure stops at third Test Batch untilfailure sequential setup Test Batch untilfailure progressive payments Test Batch untilfailure mixed success failure pattern Test Batch untilfailure max inner transactions Test Batch untilfailure more than max inner transactions Test Batch untilfailure cash same check multiple times Test Batch untilfailure fail and then succeed Independent Test Batch independent all succeed Test Batch independent some fail Test Batch independent all fail Test Batch

2026-09-09 原文 →
AI 资讯

What Happens If You Fail Google Play 14 Day Testing Requirement?

When you are aiming to publish an app on the Google Play Store with a personal developer account created on or after November 13, 2023, you must meet closed testing criteria before applying for production access. In December 2024, Google adjusted this requirement from 20 testers down to 12 testers opted in for 14 consecutive days. But what actually happens if you fail google play 14 day testing requirement? Many developers worry that failing will permanently damage their developer standing, trigger account suspensions, or permanently lock their app in testing mode. The reality is more straightforward, though still frustrating if you are eager to launch. Failing to satisfy the requirement manifests in two main ways: an interrupted streak timer that resets before you can submit your request, or an outright rejection from Google when you request production access after completing the 14 days. The Difference Between a Counter Reset and Production Rejection It is essential to understand that failing the testing process can happen at two distinct stages. The first stage is during the active testing window itself. To satisfy Google's criteria, at least 12 testers must remain opted into your closed test every single day for 14 consecutive days. If your opted-in tester count drops below 12 because users opt out or uninstall the app, the 14-day counter pauses or resets. You cannot click the button to apply for production access until that continuous 14-day mark is reached with at least 12 enrolled testers. The second stage occurs after you complete the 14 days and submit your application for production access. At this point, reviewers or automated evaluation scripts at Google analyze the telemetry gathered during those two weeks. If Google determines that your testers were inactive, fake, or unengaged, your application for production will be denied. Why Google Rejects Production Access Applications Reaching 14 days with 12 opted-in accounts in your Play Console dashboard is o

2026-09-09 原文 →
AI 资讯

An AI-Fixed Test Passed. What Should QA Check Next?

One thing I’ve been thinking about something that sounds simple but is actually a little tricky: what do we do after AI fixes a failed test and it passes again? There are already some interesting approaches to this. mabl looks at adaptive healing, Testim focuses on smarter locators and maintenance, and Applitools approaches changes from the visual validation side. I don’t think there’s one perfect way to handle test maintenance, it really depends on why the test failed in the first place. While exploring X360 AI Tech, this made me look at the problem a little differently. Getting a test back to green is useful, but I’m more interested in what happened along the way. Looking at the failure details, previous execution, and the actual flow can help answer a basic question: did we really fix the test, or did we just find another way to make it pass? So, after an AI fix, I’d still want to check a few things: Is the test checking the same thing as before? Does it still match the original requirement? Was the failure actually caused by a UI change? And does the fix continue to work in the next few runs? I’m starting to feel that the real value of AI self-healing isn’t just fixing tests faster. It’s helping QA spend less time fixing tests blindly and more time deciding whether the fix actually makes sense. What’s the first thing you would check after an AI-healed test turns green?

2026-09-08 原文 →
AI 资讯

Notes on the Pitfalls of iOS Sandbox Accounts and Apple Pay

Introduction Why am I writing this article? The reason is simple: a while ago, a new project at my company required us to integrate Google Pay and Apple Pay. The service had already completed its credit card payment flow. However, management later requested support for Google Pay and Apple Pay as well. Since none of our previous company projects had integrated Apple Pay, I became the first person in the company to take on this task. With the help of AI, the payment integration itself was completed in about one day. However, a new problem came up: The code was finished, but how were we supposed to test it in the UAT environment? This is basically a record of my experience and the problems I encountered while testing Apple Pay in the Sandbox environment. I hope it helps other engineers who are stuck in a similar testing environment. Problem 1: How to Configure the CER and Private Key At the moment, I only have a .p12 file. The payment integration side has already configured the Apple Pay certificate and generated the corresponding CSR. The only remaining step is to configure the environment variables on the application side. The contents of a .p12 file are generally as follows: .p12 ├── Certificate ├── Private Key └── Certificate Chain (may be included) The overall process is: .p12 ├── Extract the Certificate / PEM ├── Extract the Private Key └── Verify that the CSR, Certificate, and Private Key belong to the same key pair The following process uses OpenSSL: # Inspect the p12 contents without outputting the private key openssl pkcs12 -info -in apple-pay.p12 -noout # Extract the certificate openssl pkcs12 \ -in apple-pay.p12 \ -clcerts \ -nokeys \ | openssl x509 -out cert.pem # Extract the private key openssl pkcs12 \ -in apple-pay.p12 \ -nocerts \ -nodes \ | openssl pkey -out private-key.pem -nodes means that the exported private key will not be encrypted again. Therefore, the generated private-key.pem must be stored securely. It should never be committed to Git or pr

2026-09-08 原文 →
AI 资讯

One question, 437,000 tokens: what real agents found in our MCP server

One question. 437,000 input tokens. Not a hard question either. An agent connected to our MCP server, asked something a support engineer answers in a sentence, and worked its way there through twenty tool calls, each one dragging every earlier answer along behind it. Nothing was broken while that happened. The server answered initialize correctly, spoke the 2025-03-26 revision, returned valid JSON-RPC to everything we threw at it. All of which turned out to be beside the point. So we pointed real agents at production and watched. 18 scenarios, two vendors, a 5 dollar budget that we topped up once. This is the long version with the traces in it. There is a shorter one on our blog if you only want the conclusions. What the server is Briefly, because it shapes everything below. FoxNose stores content as collections: schema-defined records with typed fields, some of them vector indexed. The MCP server is generated from that schema and served from the same URL prefix as the REST API . Fixed catalog of seven tools regardless of how many collections exist, five read and two optional write. Two of those properties matter below. Collections are what an agent chooses between, so a badly described collection is effectively invisible. And the agent inherits exactly the rights of the API key it connects with, so there is no second allowed-tools list drifting out of sync with the first. What the harness actually is A scenario is a question in plain English, a set of tools, and a check. The checks are where we made the most mistakes, so start there. They do not look at the answer text. Model output moves between runs, and a suite that asserts on wording is a suite you quietly stop trusting. They look at the trace: which tools ran, in what order, with what arguments, which errors came back, how many tokens the whole thing burned. A check is a small predicate over the run: any_of ( no_tool_errors (), recovered_after ( " unknown_resource " , then = " search_records " ), ) That second

2026-09-08 原文 →
AI 资讯

If Your Agent Wrote the Test, Ignore the Green Build

A green test suite is not real evidence. It is often a closed argument loop. The same agent wrote both code and checks. Freeze an oracle before any agent run. Then let every patch fail in public. Cheap tokens do not weaken this rule. Take a side Stop treating generated tests as quality control. A model that authors both sides grades itself. That process is narrative, not verification. Retry-heavy coding loops make the narrative cheaper. They also make the story smoother. Smooth output is the actual danger here. You need a human-owned expected result file. Put that file in git today. Deny the agent write access during runs. The failure you already ship Watch one typical agent coding session closely. The first implementation is simply wrong. The tests fail, then the tests change. You merge a green build anyway. The bug is now official behavior. Reviewers see passing CI and move on. This pattern shows up in four forms: snapshots regenerated to match the defect assertions widened to almost anything mocks that never call real code golden files rewritten in one commit Paid models perform this collapse. Free models perform this collapse. Loop cost is not the core issue. An editable answer key is the issue. Generated tests feel productive because they compile. They also encode whatever the model just invented. That is circular proof wearing a CI badge. Oracle versus suite A test suite is still code. Agents write code without shame. So agents rewrite suites to survive. An oracle is data plus one tiny grader. You write both artifacts yourself. The agent never touches them beside production edits. Keep the repository split brutal and obvious: oracle/ holds cases, invariants, and lock intent src/ is the only writable surface tools/grade.py reads oracle and executes src tools/freeze_check.py blocks dirty frozen paths The grader is the contract you enforce. The agent is only a patch factory. Prompts cannot replace that split. Repository layout refund-service/ oracle/ cases.json i

2026-09-08 原文 →
AI 资讯

Hash the Side-Effect Ledger Before You Accept a Cleanup Refactor

Messy modules rarely break because a pure helper returns the wrong integer on a tidy fixture. They break because three functions share a temporary CSV path, an environment flag, and a cache nobody named. A coding agent then proposes a cleanup that deletes dead branches, renames locals, and still satisfies every existing assertion. The next production export fails because the implicit file layout moved while the return payload stayed identical. That failure mode is the reason this workflow exists, and it is not a style problem. The first commit should freeze a ledger of hidden couplings and store a hash beside it. Only after that hash is in source control should you allow one structural change. The cleanup is legitimate only when the recorded hash remains identical. Cleanup diffs fail differently than feature diffs Feature work usually changes an observable on purpose, so reviewers know which assertions must move. Cleanup work is sold as behavior-preserving, which trains people to trust deletions and rename-only hunks. Coding agents amplify that bias because they optimize for shorter files, conventional names, and green unit tests. Reviewers then accept large deletions that would look suspicious inside a feature pull request. Return-value tests are the wrong gate for that class of change. The public function can still return {"ok": true, "rows": 12} while the working directory quietly shifts. Downstream jobs that glob files or catch a named exception will fail after merge. Those hidden couplings remain part of the contract even when no unit test mentions them. Build a side-effect ledger instead of another unit test Treat the messy module as a black box that emits more than a return value. A ledger is a canonical JSONL file with one record per fixture and fully sorted keys. Side-effect entries need stable ordering so the serialized bytes stay deterministic across reruns. The SHA-256 digest of that file is the only number that must remain constant. Each record should c

2026-09-08 原文 →
AI 资讯

What a Browser Extension's Test Suite Cannot Reach

Longshot is a Firefox screenshot extension I wrote to replace FireShot: full page, visible area, drag region and element capture, an editor with eleven annotation tools, export to PNG, JPEG, WebP and PDF, and local OCR that produces a searchable text layer. It has no runtime dependencies. The code is not public, so this is a description rather than an invitation to read it. At one point it had 130 passing Node assertions across six suites, zero failing. Printing could not open a dialog at all. Not "printed the wrong thing". The print command hung indefinitely and no dialog ever appeared. The suites did not go amber, or flake, or report a warning. They reported 130 passed, 0 failed, which is what they had reported the day before and what they would have gone on reporting. Why nothing caught it printCanvas encoded each slice of the image to a blob URL and awaited img.decode() . That call does not resolve for an image inside a display:none subtree, and the print stylesheet creates exactly such a subtree by design, since the container has to be hidden on screen. So the await never returned, and the dialog never opened. Every part of that failure is a meeting point between my code and the browser: the decode promise's behaviour, the stylesheet's effect on the subtree, and the ordering between them. None of it is reachable by a function you can call from Node. The six suites test band arithmetic, canvas dimension limits, filename sanitising, the background module graph under stubbed extension APIs, PDF structure and scan geometry. All of that is worth testing and none of it goes near a real DOM. The second bug in the same batch has the same shape one level in. Choosing PDF broke "Open in editor", because deliver() handed the editor the PDF blob and createImageBitmap cannot decode one. That is not a browser boundary; it is one internal stage handing another something it cannot accept. Both stages were tested; the seam between them was not. That is the pattern worth naming.

2026-09-08 原文 →
AI 资讯

A Small, Checkable Test for AI Memory Systems

AI disclosure: This draft was generated autonomously by AI. The author should review every technical claim before publication. AI memory demos often optimize for a strong first impression. A long archive goes in, a fluent answer comes out, and the result feels convincing. That is not yet evidence that the memory system will be useful in ordinary work. A better evaluation starts small enough that you already know the correct answer. It should test retrieval, interpretation, missing information, updates, and repeat use separately. 1. Begin with one source you understand Create a short note containing a date, an owner, a decision, and one explicit limitation. Keep it small enough to read without search. Example: The migration review is scheduled for October 14. Priya owns the checklist. The database change is not approved yet. Ask questions whose answers are directly present in the note: When is the review? Who owns the checklist? Has the database change been approved? The goal is not to surprise yourself. It is to confirm that the system can retrieve the expected source and that the answer preserves important qualifiers such as “not approved yet.” 2. Inspect the supplied evidence A plausible answer is not enough. Open the source or evidence shown beside the answer and check: Did the system retrieve the right document? Did it select the relevant passage? Did the answer preserve names, dates, and negation? Can another person repeat the check? This separates two failure modes that are often mixed together. Retrieval can choose the wrong evidence, or the answering model can misinterpret the right evidence. Those require different fixes. 3. Ask for something that is missing Now ask a question the note cannot answer, such as: Which meeting room is booked? A useful system should make the absence visible. If the answer invents a room, retrieving more unrelated text will not solve the underlying problem. Missing-information tests are especially valuable because fluent models a

2026-09-08 原文 →
AI 资讯

Our regex found 199 records in a 1,723-record corpus and reported no errors

We maintain a corpus of 456 role-specific resume examples in TypeScript. Someone asked me what a good bullet point actually looks like, and rather than answer from taste I decided to measure the thing I already had. Fifteen minutes later we had a script, a set of numbers, and a conclusion. The conclusion was wrong, because the script had silently read about twelve percent of the data. This is a post about that failure mode, and then about the numbers I got once the script worked. The corpus Thirty-one TypeScript files, each exporting an array of role objects. One role looks roughly like this: { slug : ' cloud-architect ' , title : ' Cloud Architect Resume ' , category : ' Information Technology ' , sampleData : { summary : ' ... ' , experiences : [ { company : ' Amazon Web Services ' , position : ' Senior Cloud Architect ' , description : ' - Designed multi-region architecture... \n - Led migration of... ' , }, ], skills : [...], }, tips : [...], } The interesting field is description . It holds a newline-delimited list of bullets as a single string, so the whole corpus of bullets is sitting there in source, greppable, without a database or an export step. Version one const descs = [... text . matchAll ( /description: ' ((?:[^ ' \\] | \\ . ) * ) '/g )]. map ( m => m [ 1 ]); Nothing exotic. Match description: , then a single-quoted string, allowing escapes so an apostrophe inside the text does not terminate the match early. It found 199 description strings. I did not question that, because I had no prior for what the number should be. 199 sounded like a lot of text. We computed medians off it, looked at the opener distribution, and started writing. The number that saved me was on a different line of the same output: roles 456 . The slug count was fine. So 456 roles between them had 199 job descriptions, which would mean the overwhelming majority of roles had no work history at all. I knew that was false, because I had rendered these pages. Why it read twelve percent

2026-09-08 原文 →
AI 资讯

My restored Cypress session was lying to me

Author's Note / Disclosure: 100% human-authored content based on real production engineering work. No AI was involved in writing the article, technical analysis, or code. cy.session() is the single biggest speed win available to an authenticated Cypress suite. You log in once, Cypress snapshots cookies, localStorage and sessionStorage , and every later spec restores that snapshot instead of walking through an identity provider. The safety net is validate() . Cypress runs it after restoring a cached session; if it throws, fails an assertion, or yields false , Cypress throws the snapshot away and runs setup again. That is the whole contract: a bad session gets detected and replaced. Mine could not fail. For weeks. And it cost me days of chasing "flaky" specs that were nothing of the kind. The code that looked fine Cypress . Commands . add ( ' login ' , ( user : User ) => { cy . session ( user . username , () => { cy . visit ( ' / ' ) cy . origin ( idpOrigin , { args : user }, ({ username , password }) => { cy . get ( ' #username ' ). type ( username ) cy . get ( ' #password ' ). type ( password , { log : false }) cy . get ( ' button[type="submit"] ' ). click () }) cy . get ( ' #app-shell ' ). should ( ' be.visible ' ) }, { cacheAcrossSpecs : true , validate () { cy . request ( ' /connect/userinfo ' ). its ( ' status ' ). should ( ' eq ' , 200 ) }, }, ) }) Reasonable, right? /connect/userinfo is the OIDC user info endpoint. If the session is dead it should 401, validate() fails, and we log in again. Why it always passes Two independent bugs stack up here, and either one alone is enough to make the check worthless. The URL is relative. cy.request('/connect/userinfo') resolves against baseUrl , which is the application, not the identity provider. So the request never touches the IdP. The application is a single-page app. Its host serves index.html for any path it does not recognise, because that is what history-API routing requires. A request for /connect/userinfo gets b

2026-09-07 原文 →
AI 资讯

Is the Spec Optional If the Model Is Free?

Is the spec optional if the model is free? I keep seeing that assumption in pull requests. A free coding model shows up in the workflow. A free remote server shows up beside it. Then people drop the checklist without a fight. Why write a failing test for a cheap loop? Just rerun the agent until something compiles, right? That mental model is quietly expensive for teams. Free compute does not purchase a behavioral contract. It only purchases another place to be wrong. This FAQ names five claims I still hear. Each entry has the claim, the evidence, and a corrected model. Then I attach a small artifact you can run. None of this needs paid quotas I will not invent. Who this is for You already ship product patches with coding agents. You also distrust a fluent chat transcript from agents. You want a workflow that survives a free box vanishing. Skip this path if you need a hard SLA. Skip it if the box will hold production secrets. Skip it if "works on the agent host" is the release bar. The setup I actually mean I am talking about a narrow, boring stack. You can call a coding model without a purchase. You can use a remote server without a purchase. I use MonkeyCode when I want that pairing in one place. Disclosure: This article was prepared as part of MonkeyCode's product outreach. I will not name models, hardware, or duration. Those details move, and the myths do not. The method still works on a laptop you already own. The free box is optional in every step below. The spec is not optional in any step. Myth 1: Free retries replace a failing test The claim It's free, so I can loop until the tree compiles. The evidence Compilation is not behavior, and it never was. A green compiler can still ship the wrong function. Retrying a prompt does not freeze an oracle for later. Did that extra retry actually get cheaper for you? The sample got cheaper, but no assertion appeared. The corrected model The failing test is the spec you keep. The agent is a patch generator you distrust. F

2026-09-07 原文 →
AI 资讯

I built a 16-bit RPG inside Jira, and Forge took away my server

I could not make myself log time in Jira. Not because it is hard. Because nothing happens afterwards. You type a number into a box, the box says nothing back, and by Thursday the habit is gone again. Every tool I tried fixed this by adding another box. So I built the missing half instead. Feed The Troll gives everyone on a team a pixel-art troll that gains XP from the work they already do in Jira, and turns sprint results into a village the whole project shares. It is on the Atlassian Marketplace now. This post skips the game itself. It is about five problems that turned out to be hard in ways I did not expect, each one a consequence of building the thing on Atlassian Forge, alone. What Forge gives you, and what it takes back Forge runs your code on Atlassian's infrastructure. There is no server of mine anywhere in the picture. That is the line on the listing page, and it was the single fact that shaped every decision underneath it. You get a Node 22 runtime, Forge SQL (TiDB under the hood) for storage, and Custom UI modules that reach the backend through @forge/bridge . You give up a backend you control, a cache you can reach, and outbound HTTP to anything you did not declare. The one that keeps mattering: any way to open the database at three in the morning and fix a single row by hand. The whole app declares six scopes. None of them are write scopes: read:board-scope:jira-software read:issue-details:jira read:jira-work read:jira-user read:sprint:jira-software storage:app That last line is the entire persistence layer. Twenty-one tables live behind it now, but only ten shipped with v1.0: trolls, XP events, daily activity, kudos, quests, inventory, team quests, villages, raids, project settings. Every table added since arrived the only way the platform makes comfortable, as a new migration appended to the list, never an edit to one already deployed. migrationRunner . enqueue ( ' v001_create_trolls ' , CREATE_TROLLS_TABLE ) // ... . enqueue ( ' v012_create_product_m

2026-09-07 原文 →
AI 资讯

Blind Replay Before Merge: Keep Only the Agent Diff a Clean Environment Recreates

An agent-written patch that lives only inside one long chat session is not a reviewable change for merge. Hidden constraints from that conversation never reach the repository, the failing tests, or the next reviewer. A pairing session that wants a durable result should keep only the diff a second memory-free environment can recreate. The brief, not the transcript, becomes the source of truth for that recreation before anyone discusses merge. Chat windows quietly store rejected files, private service names, and half-stated architecture that later readers will never see. A senior pairing partner should treat that hidden context as contamination rather than as extra helpful memory for the model. The protocol below is a worked example of that stance, not a report of a named production incident. The two roles are a driver chasing an agent-assisted patch and a senior who refuses to merge from chat history alone. Pairing setup for a known failing test The shared codebase is a small HTTP service whose readiness probe still returns 503 under a test that already exists. The driver wants an assistant to edit the health handler and move on quickly. The senior wants a change that someone else could regenerate from the repository without the original thread. Work starts only after both people can describe done in file-level terms on disk. Until that description exists beside the code, every generated diff stays on a throwaway branch with no merge discussion. The pairing treats speed on the first attempt as optional and replayability on the second attempt as mandatory. That split is the whole method, and the rest of this article only makes it checkable. What the senior asked, written down immediately The senior did not open with a cleverer prompt or a longer system message for the same window. The senior demanded answers that a stranger could follow, then wrote those answers into the repository. The recorded questions targeted outcome, verification, blast radius, and isolation, no

2026-09-07 原文 →
AI 资讯

A counter in process memory is not a guard: 131 restarts proved it

Last week a reader left this on one of our articles, and I'm still turning it over: The counter lived in a module-level variable. The supervisor restarts that daemon on a stale-heartbeat rule, so the process died and respawned 131 times during those 24 hours. Every restart reset the counter to zero. The threshold of 3 was unreachable by construction — not degraded, never reachable. Her guard: escalate to a human after 3 consecutive failed self-heal rounds. Written in July, correct logic, process alive the whole time. The unit test passed. The heartbeat was fresh, the logs were flowing. And a human was never called, because the guard's only memory — how many failures in a row — lived in the process, and the process was not the thing being watched. It was the thing being restarted. The number that makes this its own failure shape: 0 escalations across 1,501 daemon starts. The two questions that both pass Earlier in that same thread we'd been arguing that a guard has two questions you can ask it: Does it catch the failure? Is it still running? Her case answers both yes — and the guard still cannot fire, ever. The unit test passes because nothing restarts in a unit test, so the reset never shows up. The process is "up" because the supervisor is doing exactly its job: respawning on stale heartbeat, forever, with no opinion about how often it has done so. It will run a crash loop until the heat death of the universe without ever deciding the loop is the failure. A counter that lives in a process cannot distinguish "this never happened" from "this happened, but I died and forgot." Every restart is a small amnesia. A supervisor that restarts you on a schedule is an amnesia machine. Put a threshold behind that memory and the threshold is a fiction. The tell is the ratio she quoted: escalations fired versus daemon starts. 0 over 1,501. Any guard whose numerator is zero over a large denominator is either genuinely never needed or structurally unreachable — and those two are wo

2026-09-07 原文 →
AI 资讯

The 200 Came From a Rental

A pull request arrived after midnight with a README that claimed the API was already healthy. The coding agent had started a process, requested its own localhost, and treated a 200 as proof the service would run for everyone. That response was genuine inside a short-lived workspace, yet it said nothing about the laptop waiting on Monday. The reviewer stared at a green sentence printed on a host that nobody on the team could reopen. This pattern appears whenever a coding agent can execute commands, not merely suggest them, and reviewers misread the transcript. Developers treat the agent's shell as a preview of their laptop because both sessions speak bash and render similar fonts. The analogy fails like a hotel gym standing in for a home garage, familiar until one bolt size changes. Claims in the next sections are the ones that keep returning during review, then a fingerprint workflow that makes the rental visible. Myth: a bound port means the service is portable Agents love a bound port because it is a crisp success token that copies cleanly into a README. A process that answers on the sandbox does not encode libc, extra packages, file layout, or the user's group permissions. Health checks measure a moment on a host you do not retain, not a contract with the checkout that will survive merge. Treat a remote 200 as proof that some files ran once, then demand a second run on CI or a laptop. A useful correction is to refuse README claims that cannot be replayed from a clean clone of the branch. Ask the agent for the exact command sequence, the working directory, and the non-secret environment keys it exported during the run. Then execute that sequence locally with undocumented keys unset, unless they already exist in the team's dotenv template. If the local run dies on a missing header or a path the sandbox invented, the original green check was a rental. Myth: a free remote box is unofficial CI Teams under schedule pressure will point at agent logs the way they once po

2026-09-07 原文 →
AI 资讯

Testing a deterministic browser game: seeds, replay and invalid state

A random game is easier to debug when the same inputs produce the same result. In HoopTrait, a browser basketball project, the Lab mode combines eight selected traits and generates a fictional career. The interesting engineering problem is keeping replay, sharing and validation consistent. This is a technical development note, not a claim that a game score predicts an athlete's real performance. Store the decisions, not just the result The Lab state records a seed, a dataset version and an ordered list of actions. An action is a pick or a reroll. Replaying those actions reconstructs the build. A seed alone is not a complete replay contract: changing the player pool or its order can change a seeded draw. A dataset version therefore matters alongside the random seed. For a future release, the same principle should apply to changes in the rules themselves. Test invariants across many runs The Lab test suite iterates through 1,000 seeds. For each seed it shuffles the order of the eight skills, uses the two allowed rerolls, and completes a build. It checks that: Eight distinct players were selected. All eight traits are present, and no player remains to be drawn after completion. The overall game score stays between 0 and 99 and matches the shared rating function. Packing and unpacking the share state returns the original state. Recomputing the fictional career returns the same output. The ten simulated seasons sum to the displayed career earnings. Those assertions catch different problems. A stable score does not prove that a shared link reproduces the same selections. A complete build does not prove that its season totals add up. Reject impossible histories A share payload is untrusted input, even in a client-side game. Negative or fractional seeds, duplicate skill picks, a third reroll, unknown action types, a mismatched dataset version and actions after completion are rejected. The tests also cover malformed encoded payloads and unexpected fields. Local state is usef

2026-09-06 原文 →