Lexical tokenization explained while building a lexer for a toy programming language
It's not highly theoretical and walks through actual lexer implementation in code submitted by /u/Xaneris47 [link] [留言]
找到 11105 篇相关文章
It's not highly theoretical and walks through actual lexer implementation in code submitted by /u/Xaneris47 [link] [留言]
I Spent Two Weeks Pitting Qwen 3 Max Against DeepSeek V4 I want to tell you about a rabbit hole I fell into recently. It started the way most of my projects do — someone on a Discord server I frequent asked a simple question: "Should I use Qwen 3 Max or DeepSeek V4 for my internal_compare workflow?" I had opinions, sure, but I wanted real numbers. So I cleared my calendar, fired up a couple of GPU instances, and started benchmarking. What I found surprised me, and it also reinforced something I've been saying for years: the open source ecosystem is winning, and the walled gardens of the proprietary AI world are starting to look pretty silly. Let me walk you through what I learned, the actual numbers I got, and why I keep coming back to open weight models with permissive licenses (looking at you, Apache 2.0 and MIT). Why I Care About This in the First Place I've been burned too many times by closed source vendors changing their pricing overnight, deprecating models without warning, or locking features behind enterprise tiers. You know the drill. The moment your application depends on a proprietary API, you're renting infrastructure you can't inspect, can't fork, and can't run on your own hardware. That's not a partnership — that's a leash. When a model ships under Apache or MIT, I can download the weights, audit the architecture, fine-tune it on my own data, and deploy it wherever I want. Nobody can rug-pull me. Nobody can raise prices because some quarterly earnings call didn't go their way. That's freedom, and freedom matters more than people think when you're building anything serious. So when I started this comparison, I was already rooting for the open weight contenders. But I wanted to be honest about the results, even if they complicated my bias. The Lineup I Tested Global API currently exposes 184 models through a single unified endpoint, which is honestly wild. I picked five that I thought represented the interesting tradeoffs between cost, capability, and o
Hey all! I have this thought in mind that we should all come together as a community to celebrate...
In this session, Sam Newman interviews Kief Morris and Adrian Mouat, both experts in their field. They explore the current reality of security in the container world, how infrastructure automation is impacted by latest trends, and whether platform teams are actually working. submitted by /u/goto-con [link] [留言]
After months of manual trading on Polymarket, I got tired of missing fast momentum moves on BTC, ETH,...
Fox is paying $22 billion for Roku, its streaming devices and its ecosystem.
Fox has announced that it's acquiring Roku outright, in a deal that values the streaming company at $22 billion. Once the deal is complete, Fox content will be promoted more heavily than before on Roku streamers and smart TVs. The deal will see Fox's TV networks and Tubi streamer combine with Roku's network of streaming […]
submitted by /u/parametric-ink [link] [留言]
A proposed FCC rule would kill burner phones: phones whose accounts are not attached to a particular person. The FCC plans to do this by legally forcing the country’s telecoms to store a wealth of personal information about essentially all phone customers, including a government issued identification number and their physical address, alarming privacy advocates and civil rights activists who compare the measures to those from authoritarian countries where it can be difficult to buy a mobile phone plan without giving up your identity. The proposed change would drastically shake up how people obtain phone plans in the U.S., and have all sorts of privacy and cybersecurity knock-on effects. The FCC is proposing the data collection partly as a way to combat scammers, with telecoms being required to collect other information on business and foreign customers like the intended use case of their bulk phone plan purchase and their IP address. But the changes would mean telecoms collect data on all new and renewing customers, and the FCC provides a long list of other things that the collected data could help authorities with...
Both kratom and one of its active components, 7-OH, have opioid-like effects and are widely available across the US. As health secretary RFK Jr. aims to get 7-OH banned, proponents of both are fighting.
Martin Kleppmann, an associate professor at Cambridge and author of Designing Data-Intensive Applications, discusses the evolution of data systems over the last decade, mainly the shift from monolithic databases to modular building blocks. Kleppmann underlines the importance of moving from cloud-centric data storage systems to decentralised data storage similar to Bluesky’s AT protocol. By Martin Kleppmann
In this article, the author outlines a practical approach to AI governance in the cloud, covering discovery of shadow AI, data classification at creation, IAM-based enforcement, policy-as-code, and operational controls. The article shows how organizations can embed governance into delivery pipelines, balancing security, compliance, and developer productivity without relying on manual processes. By Dave Ward
On paper, the Honor Magic V6 sounds like a tremendous leap forward for foldable phones: It's the thinnest one yet, with the biggest battery, and the best water-resistance ever. In practice, only the bigger battery feels like a meaningful improvement. The other upgrades are only fractionally superior to what came before. This isn't entirely Honor's […]
A new report warns that Miami, Kansas City, Philadelphia, Dallas, and Houston could be particularly hot places to play during the 2026 World Cup.
Check this out: i run four monetization channels side by side. Sponsored posts, display ads, YouTube ad revenue, and affiliate links. After eighteen months of tracking every dollar in a spreadsheet I built myself, I can tell you with brutal honesty: affiliate income is the only one that scales without me having to constantly produce more content or chase the next brand deal. But the math only works if you pick the right program. Most affiliates I know are promoting garbage with terrible retention, and they have no idea they're burning their audience's trust for a $9 one-time payout. Let me walk you through how I evaluate affiliate programs, what I've learned from running real funnels, and why the AI API category has quietly become the most lucrative vertical for tech creators in 2026. My Monetization Stack After 18 Months of Testing Here's a snapshot of my monthly revenue from a tech newsletter with around 34,000 subscribers and a YouTube channel sitting at 88,000 subscribers: Sponsored posts: $2,100 per placement, but I can only land maybe 2-3 per month without annoying my list Display ads: $1,800 per month from Mediavine, but this number barely moves regardless of how hard I work YouTube ad revenue: $2,400 per month, capped by watch time and RPMs Affiliate income: $6,800 per month, and it grows every single month even when I publish nothing That last number is what got my attention. Affiliate income compounds. When I published a tutorial in February recommending a tool, that single piece of content still earned me $340 in May because users stayed subscribed. No other channel behaves like that. No other channel lets a piece of content from four months ago keep paying you. But here's the catch that took me a while to figure out: not all affiliate programs are built the same way. And the difference between a good program and a bad one can be 10x in lifetime earnings per referred user. # # How I Score an Affiliate Program (The Growth Hacker Scorecard) Before I promote
If you've ever worked on a live game, you've probably watched this happen: an update ships, something feels off, the forums light up, and within an hour the community manager is in the middle of a fire they didn't start and can't immediately put out. And here's the thing most people get wrong about that person's job. "They just post updates, right?" That's the assumption. A game community manager writes patch notes, posts announcements, answers a few questions on Discord, drops the occasional meme, and keeps the social feeds warm. That's the visible 10%. The other 90% is harder, quieter, and almost invisible when it's done well. A community manager sits at a collision point. On one side: players who are frustrated because something broke, a balance change feels unfair, compensation feels insulting, or an update slipped. On the other side: a dev team that's heads-down debugging, prioritizing, and sometimes wrestling with problems that genuinely can't be fixed fast. The community manager has to talk to both sides at once — without sounding cold, defensive, fake, or corporate. That's not a "soft skill." That's translation under pressure, and it's hard. The real role: a two-way translator Strip away the memes and the role is basically two jobs sharing one desk. Job one is outward. Translate the studio to players: acknowledge the actual pain, explain what's known and unknown, hold boundaries without being defensive, and come back with real updates. Job two is inward. Translate players to the studio: take a pile of angry, contradictory, emotional posts and turn them into categorized, prioritized, actionable feedback the team can build from. Most people only ever see job one. Job two — the research half — is usually what decides whether the game actually improves. We'll get to it. Why players hate "we hear you" Players can smell a script instantly. These phrases aren't wrong , but they're hollow: "We hear you." "We value your feedback." "Please be patient." "We apologize f
Most explanations of atomic swaps stop at the spot case: two parties lock funds, one reveals a secret, both legs clear in the same short window. Clean, but it quietly assumes the trade settles right now . A lot of real agent activity isn't spot. It's a forward: two agents agree on terms today - asset pair, size, price - and settle at some future point, T+24h or T+48h. Procurement agents pre-committing to a delivery. A treasury agent locking tomorrow's FX-equivalent rate. A market-making agent quoting a forward to offload inventory risk. The economics are old; what's new is that the counterparties are anonymous software that will never meet. That raises a question spot swaps don't have to answer: what holds the trade together in the gap between agreement and settlement? In traditional markets the answer is a chain of intermediaries - a clearing house, posted margin, a credit desk that decides whether your counterparty is good for it. Strip those away, as you must in a market of anonymous agents, and the naive version of a forward collapses. If nothing binds the trade, either side can simply not show up when the price has moved against them. That's counterparty risk, and it's exactly the thing a forward is supposed to manage. This post is about how the HTLC primitive - the same hashlock plus timelock most people only associate with same-block atomic swaps - can encode a forward obligation that's binding without anyone custodying the funds in between. The timelock is doing more work than you think Recall the two parameters of a hash-time-lock contract: Hashlock: funds can only be claimed by revealing a preimage s such that hash(s) == H . The same H is used on both legs, so the act of claiming one leg reveals the secret that unlocks the other. That's what makes the swap atomic - both clear or neither does. Timelock: if the preimage isn't revealed before a deadline, the funds refund to their original owner. No third party decides this; the contract enforces it. In the sp
submitted by /u/Motor_Ordinary336 [link] [留言]
The federal government is planning to let a rule regulating federal data center operations sunset in September with no replacement.
Introduction It's crazy to me that some GitHub repos, that were just created in the last...