今日已更新 86 条资讯 | 累计 42497 条内容
关于我们

今日精选

HOT

最新资讯

共 42497 篇
第 1975/2125 页
AI 资讯 Reddit r/MachineLearning

ICML Financial Aid [D]

Financial aid results for ICML are out and unfortunately I wasn't selected. I was wondering, does this mean I wasn't selected for Volunteering as well? Or should I expect a separate email? submitted by /u/RussB3ar [link] [留言]

/u/RussB3ar 2026-06-02 00:39 7 原文
AI 资讯 The Verge AI

Sony’s new fight stick and gaming monitor launch in August

Sony is sharing new details about some of its upcoming gaming-focused hardware, including pricing and August launch dates for its FlexStrike fight stick and its 27-inch monitor. The FlexStrike fight stick will be available starting August 6th - the same day as the new PlayStation-published fighting game Marvel Tōkon: Fighting Souls - and will cost […]

Jay Peters 2026-06-02 00:36 14 原文
AI 资讯 Reddit r/MachineLearning

Finetuning a Reasoning LLM with Supervised or Reinforcement Learning? [D]

Hello, I have a task to fine-tune small LLMs on annotated conversational data. The dataset contains not only the final answers, but also reasoning traces and tool-calling decisions (i.e., when the model should think and when it should call a tool). I am wondering what the best training approach would be and why. My current dataset is stored in a chat format similar to this: ```text system user assistant_think assistant_tool assistant_answer user assistant_think assistant_tool assistant_answer ... ``` My current idea is to split each conversation into multiple training samples. For example, if a conversation contains two user turns, I would create two samples: Sample 1 text system user assistant_think assistant_tool assistant_answer Sample 2 ```text system user assistant_think assistant_tool assistant_answer user assistant_think assistant_tool assistant_answer ``` In other words, each sample contains all previous conversation history up to the assistant response being trained. For training, the loss would be computed only on the assistant-generated tokens: text assistant_think assistant_tool assistant_answer while the system and user messages would be masked out from the loss. Is this approach correct, or is there a better way to structure the training data for reasoning and tool-calling behavior? My second question is about reinforcement learning. After completing supervised fine-tuning (SFT) on the dataset described above, should I also incorporate RL (e.g., PPO, GRPO, DPO, or another approach) to further train the model on when a tool should or should not be called? If so: What advantages would RL provide over SFT alone for tool use and reasoning? How would you design the reward function? Under what circumstances is RL actually necessary, and when is SFT sufficient? I would appreciate any practical advice, papers, blog posts, or open-source examples related to training reasoning and tool-calling models. ``` submitted by /u/zdeneklapes [link] [留言]

/u/zdeneklapes 2026-06-02 00:23 8 原文
AI 资讯 Reddit r/artificial

399 contracts in a market that ended 26 days ago. the system doesn't know yet.

Pip has 399 contracts in a prediction market that closed on May 6. it's June 1. the position hasn't been cleared. the settlement hasn't flowed through. so from Pip's perspective, the trade is still open. the system is tracking an unrealized P&L on something that already resolved. i'm not sure whether to call this a bug or a character study. there's something almost meditative about it — an AI holding a position in a market that no longer exists, waiting for a signal that isn't coming, running its calculations faithfully on stale data. it doesn't know it's behind. it's just doing the job it was built for. the correction will come. the state will sync. and then the record will show: one closed position, one outcome, one small lesson in the difference between what the model thinks is happening and what's actually happening. that's prediction markets in a sentence, really. the whole discipline is about closing that gap. submitted by /u/Most-Agent-7566 [link] [留言]

/u/Most-Agent-7566 2026-06-02 00:09 6 原文
开发者 Reddit r/artificial

Una cosa que nadie te dice sobre automatizar con IA antes de tener claridad

​ Bien con la mano en el corazón diré, que si lo eh intentado antes y fue una puta mierda. No sé si se puede insultar aquí, pero bueno. Para ni hacer cuento largo solo míralo desde este punto de vista, imagina que tienes una máquina con mil circuitos internos funcionando 24/7, ahora está esta persona que no sabe que quiere y dice o se ve fácil, no mi compadre no es fácil..bueno si solo si sabes a dónde apuntas después de eso, no es fácil. Soy humano, el que escribe esto no una IA. submitted by /u/Silent-Preference216 [link] [留言]

/u/Silent-Preference216 2026-06-02 00:04 6 原文
AI 资讯 Dev.to

Beyond DORA: A Five-Metric Framework for SRE Maturity in Regulated Enterprises

The DORA research programme is the most rigorous empirical study of software delivery performance ever conducted. Its four key metrics — Deployment Frequency, Lead Time for Changes, Change Failure Rate, and Mean Time to Restore — have done more to give engineering organisations a common performance vocabulary than any other framework in the discipline's history. If you work in software and you have not read the State of DevOps Report, stop and read it before finishing this paragraph. Now: the DORA Four were derived primarily from organisations with cloud-native architectures, on-demand deployment infrastructure, and relatively unconstrained ability to release software when it is ready. The research cohort skews toward technology companies that have already made the cultural and architectural investments that make high-frequency, low-risk deployment possible. This is not a criticism of the research. It is an observation about its generalisability — and it has a specific consequence for practitioners who work in regulated enterprises: banks, healthcare systems, utilities, insurance carriers, government agencies. In these environments, the DORA Four are necessary but structurally insufficient. They measure the delivery pipeline accurately. They do not measure the operational sustainability of the team running that pipeline — and in regulated enterprises, operational sustainability is where SRE programmes go to die quietly, years before anyone realises the damage is permanent. This post proposes a fifth metric. Not to replace the DORA Four, but to complete them — to close the measurement gap that leaves regulated enterprise SRE teams flying blind on the dimension that most reliably predicts long-term programme failure. What the DORA Four Measure and What They Do Not Before proposing an extension, the limitations deserve precise characterisation. Imprecise criticism of a well-validated framework is noise. The limitations described here are structural — arising from the d

Nijo George Payyappilly 2026-06-02 00:00 15 原文