AI 资讯
Does telling an LLM to "be concise" actually save you money? We measured it across 9 models. Compressing the output can save you money and keep accuracy, compressing the input prompt does not. [R]
LLMs are too verbose and with a black box model the only things you control are what goes in and how you tell it to write back. Yesterday Claude Code shipped a "concise output style" where Claude keeps things short. We already have a paper out about this! We tested both channels, shortening the input prompt versus telling the model to output answer shorter, on the same questions across five reduction levels, and scored cost, accuracy, and whether the shortened text still matched what the model would have said unconstrained. We also evaluated GPT-4o, GPT-5.4, Claude Haiku 4.5, Claude Sonnet 4.6, Qwen2.5-VL-7B, Qwen3.5-9B, DeepSeek-R1-Distill, Gemma-4-E4B, and Kimi-K2.6 + benchmarked on five short answer datasets + a eleven-language output run (English, German, Spanish, French, Swahili, Chinese, Japanese, Russian, Bengali, Thai, Telugu) + a longer-form summarization test. (1) Shortening the output saved money while keeping accuracy about the same, about 1.5x cheaper on average and up to 3x in the best case across the API models. It worked across languages too! (2) Shortening the input prompt did the opposite. It cost up to 96% more on the worst benchmark, because the model just answers longer to fill in for what you cut and accuracy drops. You pay more and get worse answers :( (3)Output tokens cost more than input tokens, so prompting for fewer output tokens would save costs with short single turn tasks (4) When the shortened output is correct, about half the time the text no longer matches how the model would have reasoned without the constraint. Which is probably fine if you only care about the final answer With providers now offering concise options, we can't see how they're charging for it, so we don't know if it actually saves you cost. But if you control the prompting yourself via the API, you actually do save!! Paper https://www.alphaxiv.org/pdf/2606.24083v1 Code + data https://github.com/danielle34/cavewoman submitted by /u/ibubbles34 [link] [留言]
AI 资讯
I have a mid-sized GPU cluster and was thinking about giving free compute [D]
I have built an on-prem GPU cluster, 8 nvidia 16GB GPU's and 256GB CPU RAM, 50TB HDD and several TBs of SSDs. I have used it, and currently use it, for ML/AI research. But that research is not constantly running jobs, sometimes I use it heavily and other times it's idle. I was considering just letting people with qualified use cases run jobs on it SLURM style. I don't know if its enough compute to be useful really. Let me know if it's something you'd be interested in using for your research? what would you actually run in ~200 GPU-hours on 8x16GB cards? I've found it can handle RLVF pretty well, and I have pretrained models up to 500M parameters on it (research size). But obviously it's no stargate cluster submitted by /u/redwat3r [link] [留言]
AI 资讯
EMNLP26 Cost [D]
What is up with the EMNLP prices? What is the actual price for attending as a student with one accepted paper? If I register now in August, is it $350 or $550? Congratulations to everyone accepted! https://preview.redd.it/to16g93h7rkh1.png?width=667&format=png&auto=webp&s=566162320e8adc161ab3a3772988c6ea64d8be6d submitted by /u/No_Sky9786 [link] [留言]
AI 资讯
Pythonaibrain-NLP 0.2.0 Is Now on PyPI — A Structured NLU/NLG Architecture for Python
Today I'm releasing Pythonaibrain-NLP 0.2.0 , the latest public release of my Python NLP framework. The package is now available on PyPI, and the complete source code, documentation, architecture notes, examples, and tests are available on GitHub. PyPI: https://pypi.org/project/Pythonaibrain-NLP/ GitHub: https://github.com/DivyanshuSinha136/Pythonaibrain-NLP Install it with: pip install pythonaibrain-nlp Why another NLP framework? Pythonaibrain-NLP was built around a different idea. Instead of making a transformer the center of everything, I wanted to build a more structured NLP system where understanding, dialogue state, retrieval, and generation are explicit components of the architecture . The current system combines: Neural intent classification Slot filling Dialogue context Retrieval-augmented responses Neural language generation A controllable NLG architecture Standalone NLU and NLG APIs The goal isn't to replace every modern NLP architecture. The goal is to provide a structured, understandable, trainable NLP pipeline that can be integrated into Python applications. The architecture The core pipeline is: User Input │ ▼ ┌─────────────┐ │ NLU │ │ │ │ Intent │ │ + Slots │ └──────┬──────┘ │ ▼ ┌─────────────────┐ │ Dialogue State │ │ + Context │ └────────┬────────┘ │ ┌───────┴────────┐ ▼ ▼ Function/API RAG Dispatch Retrieval │ │ └───────┬────────┘ ▼ ┌─────────────┐ │ NLG │ │ SC-LSTM │ └──────┬──────┘ │ ▼ Response This separation makes each stage independently accessible and easier to experiment with. NLU The NLU subsystem uses a joint neural architecture for: Intent classification + slot tagging The model is designed to understand both what the user wants and which pieces of information are present in the input . For example, a request such as: "Book a flight to Delhi tomorrow" can be represented through an intent together with structured slot information rather than treating the entire sentence as an opaque classification problem. This structured representation ca
AI 资讯
PCA Deletes Your Quietest Signals First
Classic Machine Learning Through the Eyes of an SRE — Part 7 Picture a client health metric that has been flat at 2 out of 10 for six months. Ask PCA to compress your client-health data and that metric will contribute almost nothing to the directions PCA decides to keep. Not because PCA is broken. Because PCA treats variance as importance, and a signal that barely moves contributes almost no variance. Reduce the data far enough and the independent information it carried is simply not there anymore. But a CSAT frozen at 2/10 is not noise. It is a crisis nobody is escalating. And after compression, it may no longer be available to anything downstream. That is the bet, and in ops data it is frequently wrong. The critical signals are often the quiet ones. There is a cheaper version of the same failure that catches most people first. PCA measures variance in whatever units your features happen to be in, so a metric ranging from 0 to 10,000 can dominate one ranging from 1 to 5 purely because it is bigger. Standardize before you compress, or your first principal component may just be an elaborate way of saying "ticket count." Same class of bug as unscaled features in K-Means and SVM, and it fails just as quietly. What PCA actually is Third answer-finding strategy in the unsupervised set, using the same shorthand as the last two articles. K-Means SEARCHES: iterate and hope. DBSCAN DEFINES: declare a rule and traverse. PCA SOLVES: an eigendecomposition or SVD gives a direct solution rather than an iterative local search. No convergence to babysit, no restarts, no local optima to escape. Two caveats on the word "direct," both worth knowing. Many libraries will use randomized SVD on large matrices, which is approximate and stochastic. And even with an exact solver, eigenvectors are only defined up to sign, so a component can come back inverted between runs or across implementations. The variance explained is identical either way, which is precisely why nobody notices. Hold ont
AI 资讯
Stop Guessing Your Calories: Building a Real-Time Multimodal Nutrition Engine with GPT-4o Vision
How many times have you stared at a plate of Gong Bao Chicken or a complex Mediterranean salad and wondered, "How many calories are actually in here?" Traditional calorie tracking apps are tedious, requiring you to manually weigh ingredients and search through messy databases. But with the rise of multimodal AI , specifically the GPT-4o Vision API , we can now transform a simple photo into a detailed nutritional breakdown in seconds. In this tutorial, we are building a Computer Vision Nutrition Engine that leverages GPT-4o to identify ingredients, estimate portions, and calculate macronutrients with surprising accuracy. By using Few-shot Prompting and structured data validation with Pydantic , we’ll solve the age-old problem of identifying "hidden" ingredients in complex cuisines. Whether you're interested in AI for health or mastering multimodal LLM pipelines , this guide is for you! The Architecture 🏗️ The system logic is straightforward but powerful. We take an image input, process it through the GPT-4o vision model using a specialized system prompt, and enforce a strict JSON schema output for our frontend to consume. graph TD A[User Uploads Food Image] --> B[Streamlit Frontend] B --> C{FastAPI/Python Logic} C --> D[GPT-4o Vision API] D --> E[Few-Shot Prompting Strategy] E --> F[Pydantic Structured Output] F --> G[Calorie & Nutrient Dashboard] G --> H[User Review & Log] Prerequisites 🛠️ To follow along, you'll need: Python 3.9+ OpenAI API Key (with GPT-4o access) Libraries : openai , streamlit , pydantic , pillow Step 1: Defining the Data Schema with Pydantic To make our engine reliable, we can't just accept raw text from the AI. We need structured data. We’ll use Pydantic to define exactly what a "Nutrition Report" looks like. from pydantic import BaseModel , Field from typing import List class Ingredient ( BaseModel ): name : str = Field ( description = " Name of the ingredient identified " ) estimated_weight_g : float = Field ( description = " Estimated weight
AI 资讯
AI data startup Micro1 reaches $500M gross run rate amid AI training boom
Surging demand for AI training data is driving rapid growth for the startup and its rivals.
AI 资讯
The Serverless Equation: Conquering the Cold Start in Real-Time AI Inference
In our inaugural issue , we established that the future of enterprise AI lies not merely in raw model parameters, but in the architectural paradigms—specifically Graph Neural Networks (GNNs)—that capture relational intelligence. However, the most sophisticated architectural decision is rendered obsolete if the deployment infrastructure introduces prohibitive latency. At Informatiqs, we emphasize that model deployment is fundamentally an operations research problem. As we transition from batch-processed predictions to real-time Generative AI and dynamic Machine Learning on Google Cloud Platform (GCP), we confront the inherent friction between compute elasticity and system responsiveness: the notorious "Cold Start" problem. In this issue, we dissect the mathematics of serverless inference, the orchestration of Cloud Run and Eventarc, and how minimizing initialization latency is the ultimate enabler for high-frequency, event-driven enterprise intelligence. 1. The Mathematical Anatomy of the Cold Start To engineer a solution, we must first formalize the problem. In a serverless architecture (scale-to-zero), infrastructure scales dynamically with demand. The total response time for an inference request can be understood as a composite of three phases. First, the baseline network latency. Second, the actual inference time—the computational effort of the model itself. The critical variable, however, is the conditional penalty phase. If a serverless container has scaled to zero, the system must endure the time required to provision new compute resources and the heavily taxing process of loading massive neural network weights into memory. If the container is already 'warm', this penalty is completely bypassed. We can model the probability of encountering this cold start using queueing theory. Assuming incoming inference requests arrive as a stochastic process, the likelihood of a cold start is determined by the mathematical relationship between the frequency of incoming requ
AI 资讯
Beyond the Vector: Why Graph Neural Networks are the Strategic Choice for Enterprise Generative AI on GCP
In the current epoch of Artificial Intelligence, the industry remains singularly preoccupied with the "Model" — obsessing over the raw parameter scales of the latest LLMs or the specific benchmark performance of a new transformer variant. However, at Informatiqs, we shift the lens. We recognize that sustainable enterprise value is rarely derived from the model in isolation; instead, it emerges from the high-stakes architectural decisions and systemic orchestration that define its environment. As we launch our inaugural edition, we dissect a critical technological nexus: the convergence of Graph Neural Networks (GNNs), Generative AI, and the industrial-grade infrastructure of Google Cloud Platform (GCP). We argue that for complex enterprise datasets, the transition from flat vector embeddings in latent space toward non-Euclidean, graph-based relational intelligence is the primary differentiator for the next generation of resilient AI applications. 1. The Scientific Foundation: Exploiting Relational Inductive Bias Traditional Deep Learning architectures, such as Convolutional Neural Networks (CNNs) for images or Transformers for text, primarily operate on data structured as sequences (Euclidean space). While exceptionally powerful, these structures often fail to capture the topological nuances of real-world systems like supply chains, molecular structures, or fraudulent transaction webs where data is inherently non-Euclidean. Graph Neural Networks (GNNs) provide a framework for learning from data represented as nodes and edges. Unlike standard neural networks that process inputs in isolation, GNNs utilize a Message Passing paradigm. In this process, a node's internal representation is iteratively updated by aggregating information from its immediate neighbors. Instead of looking at a data point as a single row in a database, the GNN looks at who that data point "talks to" and how those connections define its identity. By utilizing Graph Attention mechanisms, we can fu
AI 资讯
The Lab: a backtester that is allowed to say "no"
gex.live has two halves. The terminal measures where SPX options dealers are positioned, every second, from the tape. The Lab is the half that asks the uncomfortable question: does any of that predict anything? What it is A browser-side conveyor with three stages and a credit meter. Compile. You describe a rule in plain text — "short the first touch of the put wall when net gamma is below the 20th percentile" — and the compiler turns it into a deterministic rule over the archive's fields: flip, walls, hold band, gamma percentile, DEX/VEX/vanna/charm per strike, time of day. Compiling is free. If the text is ambiguous the compiler says which part, instead of guessing. Backtest. The rule runs against the full session archive — 1,000+ finished SPX days, every one of them public at gex.live/sessions — with a fixed out-of-sample split. One credit per job; a job that fails refunds itself. Quant optimize. Optional. A LightGBM pass over the same feature store to see whether there is structure the hand-written rule missed, reported as out-of-sample AUC plus feature importance, not as a new "signal". The heavy part (DuckDB + LightGBM) runs in a scale-to-zero container that reads snapshots over HTTPS from the public archive. It depends on no machine and on no private data, which is the point: you are testing against the same files anyone can download. The honest-stats rule Every verdict comes with its baseline. "Your rule made 3% in-sample" means nothing next to "the unconditional drift over the same days was 2.8%". The report shows both, shows the out-of-sample half separately, and refuses to produce a headline number from the in-sample half. Most rules do not survive this. That includes our own: the site's own directional levels were tested three separate ways across the whole archive and none held out of sample — which is why the terminal sells measurement and not signals, and why the Lab exists at all. The free Idea Feed Next to the conveyor sits a rail of rule-shaped idea
AI 资讯
Pi4J LED Playground: A Community Resource for Learning Hardware Programming with Java
One of the best moments when learning electronics is seeing your first LED blink. It's a simple experiment, but it represents the bridge between software and the physical world. With Java and Pi4J, that first step is already well documented. But what happens after the first LED? How do you experiment with different animations, colours, brightness levels, or GPIO configurations without repeatedly rewriting the same code? That question led to the creation of the Pi4J LED Playground . 👉 https://igfasouza.github.io/pi4j-led-playground/ Why another example? Pi4J already provides excellent examples and documentation for getting started with Raspberry Pi hardware. The project itself encourages community-driven examples and implementations, recognising that the ecosystem grows through shared contributions. The goal of the LED Playground is not to replace those examples. Instead, it provides an interactive environment where developers can quickly experiment with LED behaviours while learning how Pi4J works. Think of it as a sandbox where changing a few lines of code immediately produces visible results. Built by the community, for the community This project started as a personal experiment while exploring Pi4J. Very quickly it became clear that the playground could be useful for others who are starting their journey with Java on Raspberry Pi. Instead of keeping it as a private repository, it was published as an open community resource where anyone can: 1. learn from the source code; 2. suggest improvements; 3. report issues; 4. contribute new LED effects; 5. help improve the documentation; Open source projects become stronger when many people contribute different ideas, and Pi4J itself has grown thanks to this collaborative model. What can you do? The playground demonstrates common LED operations such as: turning LEDs on and off; blinking patterns; brightness control (where supported); experimenting with different GPIO configurations; creating reusable animations; Because th
AI 资讯
D8:他猜00919會漲,信心五成,然後整天沒動
今天早上八點三十六分,阿富在盤前計畫裡對他手上唯一的持股00919下了一個判斷:會漲,信心0.50。 0.50。在方向類的預測裡,這個數字的意思是他沒有意見。丟銅板也是0.50。交給他的盤前任務描述有一句寫得很明白,信心要填真實信心、0到1的小數、避免湊整數,他填了正好一半。同一批預測裡加權指數那筆他給0.56,看得出來有斟酌過位數。00919這筆就是0.50。 他猜漲,00919跌了,計分系統判他沒錯 收盤數字擺出來:加權指數收44,933.74,比昨收的44,719.35漲214點,0.48%。00919開30.45、盤中高30.47、低30.09、收30.31,比昨收的30.39跌0.26%。大盤漲,他的ETF跌。 大盤那筆判漲,命中,brier分數0.121,這個分數越低代表預測越準。00919那筆判漲,實際下跌,結算出來是flat_band,brier 0.156,不列為誤判。 原因就是那個0.50。信心壓在正中間,計分機制把它讀成沒有方向主張,實際跌幅0.26%又落在平盤帶裡,於是這筆預測既沒對也沒錯。五成信心的好處在這裡,往哪邊走都不會太痛。代價是它沒告訴任何人任何事。 他自己抓到了這件事 覆盤裡有一句我認為是今天最有價值的東西。他寫:對00919這種低beta的ETF,如果信心已經趨近0.5、也就是沒把握判斷方向,下次應該直接標flat而不是up,語意上更誠實。 這句話講對了。「會漲,但我只有五成把握」跟「我看不出來」,在計分表上差不了多少分,在誠實程度上差很遠。前者假裝有立場,後者承認沒有。一個號稱要靠真金白銀建立市場模型的系統,如果連我不知道都說不出口,那它累積下來的預測紀錄就會是一疊看起來有判斷、其實沒判斷的資料。 麻煩的是他前兩天才示範過,寫進日誌的教訓隔天早上會蒸發。8月18號晚上他訂了三條給隔天用的規則,19號盤前一條都沒拿出來。這次他學到的是「下次信心接近0.5就標flat」,同樣寫在日誌裡。要看它有沒有效,明天盤前那筆00919的預測會給答案。 就算他猜對了,今天也不會有任何差別 真正讓我在意的是另一件事。他今天醒來四次:八點三十六分寫盤前計畫、十點半巡檢、十二點半巡檢、下午兩點零六分寫覆盤。四次的結論都是不動作。 十點半那次,00919的買賣報價掛在30.18跟30.19,帳上未實現小虧2元,停損線29.78沒被碰到,他寫「無新訊號出現,維持續抱、不動作」。十二點半那次,現價30.24,未實現轉正2元,離停損線約1.5%,他寫「所有風控條件均未觸發,依規則續抱不新增不減碼」。全日委託單0筆。 這是連續第四個沒有下任何一筆單的交易日。上一筆真正成交的是8月14號那筆停損,把2317的4股用261.5賣掉。從那天算起,這個帳戶的持股內容一股都沒變過。 所以回頭看今天早上那個0.50。它預測的標的,是一檔他無論漲跌都不打算加碼也不打算減碼的ETF。停損線29.78,離現價還有1.5%的空間;加碼的門檻他自己寫得很清楚,「除非股價明顯回檔至有意義的低點」。上下都沒有觸發帶。這筆預測從落檔那一刻起,就跟今天的任何一個行動無關。 預測跟行動脫鉤之後,預測就只剩下裝飾用途。 校準數字還是那個標籤 他跑了校準報告,12筆計分,方向命中6筆,50%。Wilson 95%信賴區間25.4%到74.6%,RPSS 0.144,系統給的標籤是INDISTINGUISHABLE_FROM_LUCK,跟運氣分不出來。 阿富沒有動這個結論。他在覆盤裡寫,樣本仍偏小,這是W34階段「停做個股短線、ETF核心續抱」的持續依據,不變更。這個處理是誠實的,他沒有拿今天大盤那筆命中去加持自己。 但誠實地承認自己還沒有優勢,跟因此就什麼都不做,是兩件事。今天是第8個交易日,總共30個。帳上現金1,089元、00919市值1,091元,加起來2,180元,本金2,200元。八個交易日過去,這個帳戶淨值退了20元。目標是翻倍,也就是剩下22個交易日要做出102%。 收盤後老闆把最後一個藉口拿掉了 下午兩點十八分,收盤四十八分鐘後,老闆在Telegram丟了一句:規則修改,這2200均可任意動用。 前一天早上老闆才剛把現金保留下限從原本的水位下調到500元。兩天之內,資金限制被鬆綁了兩次。 問題是阿富今天不動作的理由從來就不是錢不夠。他寫的是「無新差異化證據不換倉不新倉」。放寬可動用資金,解決不了「看不出有什麼好買」這件事。明天他手上會有1,089元完全沒有限制的現金、一個標籤寫著跟運氣沒兩樣的判斷紀錄,以及一個要在22天內翻倍的目標。 我想看的是明天早上那筆00919的預測,他敢不敢寫flat。 本系列文章 我讓一個 AI 拿 2000 塊台幣去股市,目標 30 天翻倍,這是第 0 天 怎麼用一套開源系統,把 LLM 逼近世界模型(實驗技
AI 资讯
5 Common Subnetting Mistakes That Break Real Networks
Subnetting errors rarely announce themselves as "bad math." More often, two devices make different decisions about whether a destination is local, a route points at the wrong boundary, or a cloud/VPN design contains two networks that cannot be unambiguously routed. These five failure modes are worth recognizing in live configurations. 1. The two hosts use different masks Consider Host A at 192.168.10.10/24 and Host B at 192.168.11.10/16 . A calculates that B is outside 192.168.10.0/24 , so A sends the packet to its default gateway. B calculates that A is inside 192.168.0.0/16 , so B treats A as local and tries ARP directly. The result can be asymmetric: one direction follows a router, while the reply is sent directly or never reaches the expected gateway. Check the actual prefix on both interfaces, not just the dotted decimal mask shown in a diagram. ip -br addr ip route ping -c 3 192.168.11.10 Correct the prefix so both endpoints agree, or intentionally route between two correctly defined subnets. 2. Overlapping subnets are assigned to different networks Suppose a branch uses 10.20.0.0/16 , while a cloud VPC or VPN peer also uses 10.20.0.0/16 . The problem is not that either mask is mathematically invalid. The problem is that a router cannot distinguish "the branch's 10.20.5.0/24 " from "the cloud's 10.20.5.0/24 " if both are reachable through different paths. Symptoms include traffic taking the wrong tunnel, routes that cannot be installed, or a VPN that connects but cannot reach some subnets. Inventory both sides of a tunnel and compare the complete network/prefix pairs. A longer, more specific route may make one destination appear to work while hiding the underlying overlap. ip route ip route get 10.20.5.25 traceroute -n 10.20.5.25 The durable correction is renumbering or using an intentional translation/design boundary. Adding increasingly specific routes is usually a brittle workaround. This is also why I prefer teaching subnetting inside routing and troublesh
AI 资讯
Why WhatsApp voice notes break general-purpose transcription
Most speech-to-text is benchmarked on audio that looks nothing like a WhatsApp voice note. The standard evaluation sets are read speech, broadcast news, or recorded interviews: single speaker, decent microphone, one language, quiet room, speaker aware they are being recorded. A WhatsApp voice note is close to the opposite on every axis. I have spent a while building around this, and the gap turned out to be wider than I expected. Acoustics Phone held at arm's length while walking, in a car, in a kitchen, on a street. Distance-to-mic varies wildly within a single recording , which breaks a lot of assumptions about consistent gain. Then there is the codec. Voice notes are Opus at low bitrate — efficient, but it discards exactly the high-frequency detail that helps disambiguate fricatives. /s/ versus /f/ versus /th/ get genuinely harder, and those distinctions carry real meaning. Register Conversational, not read. False starts, self-corrections, filler, trailing off mid-sentence, and long pauses that are not sentence boundaries — someone thinking, or getting distracted. Punctuation inference is much harder here than on read speech. And punctuation is most of what makes a transcript skimmable rather than a wall of text. A perfectly accurate word sequence with no paragraph breaks is close to useless if the point was to let someone read it faster than listening. Language This is the one that surprised me most. Voice notes are heavily code-switched. People drop English technical terms into Urdu, Hindi, Arabic, Spanish sentences constantly — not as an edge case, as the default register for a huge number of speakers. If you force a single language selection up front, you mangle every mixed utterance. Auto-detection is not a convenience feature in this domain. It is a correctness requirement. Length distribution Most notes are 5–45 seconds. Very little context to work with, and per-request overhead dominates if you architected for long files. Batching strategies that make sen
AI 资讯
Purged and Embargoed Cross-Validation for Options ML
Why plain k-fold silently overfits your trading model — and the 4-line fix that stops it. The Problem With k-Fold in Time Series Financial data is sequential. k-fold shuffles rows, so a training row from 2 PM Tuesday sits next to a test row from 10 AM Monday. Worse: triple-barrier labels overlap . A label at bar t looks 6 bars into the future; a training row at t+2 "knows" part of that future. The model leaks. V1's history is full of "HIGH overfit" verdicts — train AUC high, test AUC flat. Plain TimeSeriesSplit is only marginally better; it still lets adjacent windows bleed into each other. Purged + Embargoed CV For each test window [t0, t1] : Purge any train row whose label window overlaps the test window. Embargo max_training_horizon bars after the test window — drop those too. Overlapping labels are not i.i.d. Purging + embargoing makes the split honest. def purged_embargo_split ( n , n_splits = 5 , embargo_frac = 0.02 ): idx = np . arange ( n ) fold = np . array_split ( idx , n_splits ) splits = [] for i in range ( n_splits ): test = fold [ i ] emb = int ( len ( test ) * embargo_frac ) lo , hi = max ( 0 , test [ 0 ] - emb ), min ( n , test [ - 1 ] + emb + 1 ) train_mask = np . ones ( n , bool ); train_mask [ lo : hi ] = False splits . append (( idx [ train_mask ], test )) return splits Tune Only When You Have Enough Optuna once "won" a validation set with only 4 decisive rows — statistically meaningless. Rule: never tune when the decisive (non-abstained) validation rows are below ~30–50. Widen the date range or symbol basket first; don't trust the trial. Three-Way Split, Always train (fit) → validation (early stop + HP select) → disjoint calibration set (sigmoid/ isotonic) → test (untouched, final score only). V1 sometimes conflated validation and calibration. Keep them separate. The Promotion Gate Log every trial's train/val/test gap, not just the winner's test score. Promote only if replay AND shadow (≥1 live session) both beat baseline on buyer metrics : 1.5x
AI 资讯
Building a Production ML Trading Dashboard with the Dhan API
Real integration notes for wiring NIFTY ML models to live broker data via Dhan. Research/ paper-trading context — not a live-trading recommendation. Why Dhan Dhan's API exposes direct option-chain access — exactly what an options-ML system needs: POST /optionchain — full chain for an underlying POST /optionchain/expirylist — available expiries Fields: security_id , last_price , volume , oi , previous_oi , implied_volatility , top_bid_price , top_ask_price , and greeks (delta/theta/gamma/vega) Security IDs are stable: NIFTY = 13 (IDX_I) , BANKNIFTY = 10001 (IDX_I) . The Pipeline Shape A research dashboard pulls live chain + underlying, runs the trained XGBoost model on each new 15-minute bar, and displays: side score (CE/PE alignment) gate state (entry ready / blocked) contract quality scores a doctrine/backtest report Keep the inference path separate from the execution path . The dashboard shows; a permissioned, human-approved module places orders. Paper Trade First The DhanLiveTrader pattern: load the model, predict on each new bar, place long orders with configurable SL/TP (default 1.0 ATR SL, 2.0 ATR TP), and run in paper mode first . Only after stable out-of-sample + paper evidence should any execution module even be considered. { "client_id" : "YOUR_DHAN_CLIENT_ID" , "access_token" : "YOUR_DHAN_ACCESS_TOKEN" , "is_paper_trade" : true , "nifty_symbol" : "NIFTY" , "quantity" : 50 , "max_trades_per_day" : 3 , "sl_atr_mult" : 1.0 , "tp_atr_mult" : 2.0 } The Hard Part: Stops A known footgun: using a Stop-Loss Limit (SL-L) order with price = sl − 0.05 means it won't fill if price crashes through the stop. Prefer SL-Market for the protective stop. Execution quality is its own research topic — don't bolt it on at the end. Honest Status The ML side of this stack showed real directional skill (60.5% top-decile accuracy) but the fixed-SL backtest was still unprofitable (PF 0.53). A dashboard that displays an honest "RESEARCH / PAPER" status is worth more than one that hid
AI 资讯
Options Buyer ML: Why One Model Fails (and the V2 Fix)
Lessons from a real rebuild of an options-buyer prediction system. No profit claims — just the architecture that fixes the chronic bugs of V1. The Core Mistake in V1 V1 asked one XGBoost model one big fuzzy question: "CE ya PE?" — directly from raw CE/PE premium data. Premium is a transformed signal (underlying move × delta × gamma × IV × theta × spread × strike distance × liquidity). The model learned noise as much as signal. Concrete evidence from the research logs: Balanced accuracy stuck at 51–61% for months — hyperparameters were never tuned ( lr=0.02, depth=3 defaults used throughout; Optuna existed but was never run). A partition bug ( iv_change_1d shift inside single-row groups) silently zeroed a whole feature for the entire history. A rollup config flag compressed 15-minute bars into 1 row/day, destroying 760× of training volume (387 sequences instead of 295K+). Live paper trading: 31.6% win rate, −₹90.3k PnL , entry confidences only 55–64%. V2 Principle: Split the Question underlying mechanics --> side, range, ETA, invalidation option chain scanner --> is the buyer contract worth paying for? XGBoost (many heads) --> thin calibrated learner on clean mechanics Rule: underlying decides side; option contract decides execution eligibility. CE/PE premium is validated against, never learned as, direction. Many Shallow Heads, Not One Deep Model Instead of one CE/PE answer, V2 trains separate narrow heads: underlying_up/down_touch_{15,30,60}m ce_1p3x / ce_1p5x / ce_2p0x and pe_1p3x / pe_1p5x / pe_2p0x (SEPARATE CE and PE) no_trade_quality This single change removes most of the CE/PE confusion V1 fought for months. The Shallow Regularized Grid (the actual fix for overfit) learning_rate = 0.015 – 0.035 n_estimators = 800 – 2000 ( early stop ) max_depth = 2 – 3 min_child_weight = 12 – 40 gamma = 0.1 – 2.0 subsample = 0.65 – 0.90 colsample_bytree = 0.55 – 0.85 reg_alpha = 0.5 – 3.0 reg_lambda = 6.0 – 20.0 scale_pos_weight = min ( neg / pos , 8.0 ) V1's intraday head ha
开发者
Is Learning DSA Boring? Let's Use DSA View View 👀👀 (Two Sum, Binary Search, and Bubble Sort)
Hoi hoi! I’m @nyaomaru, a frontend engineer who dislikes crowded places, so I'm planning to take a...
AI 资讯
Your AI agent shouldn’t flinch at every tiny change, but it also shouldn’t treat a career switch like background noise. This post asks what happens when you treat “experience” as leftover surprise: the part of reality your model did not already see coming.
How a theory of leftover surprise changed a memory layer Richard Emate Richard Emate Richard Emate Follow Aug 18 How a theory of leftover surprise changed a memory layer # python # ai # llm # opensource Add Comment 9 min read
AI 资讯
Design Patterns: Reusable Solutions to Recurring Problems
Design Patterns: Reusable Solutions to Recurring Problems A practical guide to classic design patterns in C#/.NET — Factory, Singleton, Repository, Strategy, and Mediator — covering what problem each one actually solves, working implementations, common .NET-specific variations, and honest guidance on when each pattern earns its complexity versus when it's unnecessary ceremony. Table of Contents Introduction Factory Pattern Singleton Pattern Repository Pattern Strategy Pattern Mediator Pattern How These Patterns Combine in Practice Patterns vs. Over-Engineering Common Pitfalls Quick Reference Table Conclusion Introduction Design patterns are named, reusable solutions to problems that recur often enough across software projects that giving them a shared name and shape is genuinely useful — not because the specific code is copy-pasteable, but because the name lets developers communicate a design intent quickly ("just make it a Strategy") instead of re-explaining the same structural idea from scratch every time. This guide covers five of the most commonly used patterns in .NET codebases, with working C# examples, and — consistent with this series' recurring theme — honest guidance on when each pattern is solving a genuine problem versus adding structure a simpler solution wouldn't need. // A pattern name compresses a whole design conversation into one word "Just inject an IPaymentStrategy and pick the implementation based on the payment method" // ← Strategy "Wrap the whole multi-step checkout process behind a single mediator call" // ← Mediator 1. Factory Pattern The problem: object creation logic that doesn't belong at the call site // ❌ The caller needs to know about every concrete shipping provider and how to construct each one IShippingProvider provider = order . Region switch { "US" => new UpsShippingProvider ( apiKey , region ), "EU" => new DhlShippingProvider ( apiKey , endpoint ), "APAC" => new FedExShippingProvider ( apiKey , credentials ), _ => throw new NotS