今日已更新 189 条资讯 | 累计 41866 条内容
关于我们

标签:#p

找到 17352 篇相关文章

AI 资讯

Notes on Federated Learning and Differential Privacy

Notes on Federated Learning and Differential Privacy 2026-05-31 · privacy-preserving ML Working notes on building federated learning (FL) from scratch, what actually breaks under Non-IID data, and how differential privacy (DP) and secure aggregation fit on top — including the honest negative results that the marketing slides leave out. They follow the implementation in federated-learning-lab (FedAvg / FedProx / SCAFFOLD, DP-SGD, secure aggregation; 33/33 tests, literature cross-validated). 1. What federated learning actually is The data never moves. Instead of pooling everyone's data on one server, each client trains locally and sends model updates to a server that aggregates them. The canonical loop ( FedAvg ) is: Server broadcasts the global model. Each client does a few local SGD epochs on its own data. Each client sends back its updated weights. Server averages the weights (weighted by client data size) → new global model. That's it. The elegance is that raw data stays on-device; the difficulty is that the clients' data distributions are not identical. 2. The Non-IID problem (where FedAvg starts to hurt) FedAvg implicitly assumes every client sees roughly the same distribution. Real clients don't — one hospital sees different cases than another, one phone's keyboard sees different language. Under Non-IID data, each client's local optimum pulls in a different direction, so averaging their updates produces client drift : the global model lands somewhere none of them wanted. Two well-known fixes, both implemented and measured in the lab: FedProx — add a proximal term that penalises drifting too far from the global model. Stabilises training when clients are heterogeneous. SCAFFOLD — track control variates (correction terms) that estimate and subtract the drift direction. More state to communicate, but corrects the bias FedProx only damps. The honest finding worth repeating: on a strongly Non-IID split (e.g. label-skewed MNIST), the fancy methods don't always beat p

2026-05-31 原文 →
AI 资讯

Notes on Serving LLMs with TensorRT-LLM and Triton

Notes on Serving LLMs with TensorRT-LLM and Triton 2026-05-31 · LLM serving / NVIDIA stack These are working notes on taking an open-weights LLM from a Hugging Face checkpoint to a production-style serving endpoint on the NVIDIA stack — TensorRT-LLM for the engine, Triton Inference Server for the deployment surface — and benchmarking it honestly against vLLM on multi-GPU hardware. They follow the harness in trtllm-triton-serving (4× H100, NVLink). The goal is to move from "I use vLLM" to "I can stand up the NVIDIA inference stack on real multi-GPU hardware and reason about the trade-offs." 1. The serving pipeline The path from checkpoint to endpoint has four stages. Each one is a place where a decision affects latency, throughput, or accuracy: Checkpoint — a Hugging Face model. Engine build — compile to a TensorRT-LLM engine for a fixed tensor-parallel degree, precision, and batching policy. Model repository — wrap the engine in a Triton tensorrt_llm -backend model repo. Serving + load test — trtllm-serve (or Triton) exposes an OpenAI-compatible endpoint; a load generator drives it under controlled concurrency. The key mental shift from vLLM: TensorRT-LLM does ahead-of-time compilation . vLLM is a runtime that takes the model and serves it; TensorRT-LLM builds an engine specialized to your GPU, TP degree, and precision first. That build is where the performance comes from, and also where the rigidity comes from. 2. Tensor parallelism (TP) For a model that doesn't fit on one GPU — or to cut latency — TensorRT-LLM shards each layer across GPUs. On a 4× H100 NVLink box, TP=4 means every forward pass does an all-reduce across the four GPUs over NVLink. The all-reduce is not free. On this fabric it tops out around 77 % of the NVLink budget (see the separate NVLink-wall notes ). For prefill (large tensors) you're bandwidth-bound and TP helps. For decode (one token at a time) you're pinned against the small-message latency floor, and past a point more TP makes decode slowe

2026-05-31 原文 →
AI 资讯

JWT Explained: What's Actually Inside That Token (with a free decoder)

If you've ever worked with auth, you've seen a JWT — a long string like eyJhbGci... split into three parts by dots. It looks cryptic, but it's surprisingly simple once you see inside. A JWT has three parts header.payload.signature Header – tells you the signing algorithm, e.g. {"alg":"HS256","typ":"JWT"} Payload – the claims (the actual data), e.g. {"sub":"123","name":"John","iat":1516239022} Signature – verifies the token wasn't tampered with (needs the secret/key) The header and payload are just Base64url-encoded JSON — not encrypted. That means anyone can read them. So: ⚠️ Never put secrets in a JWT payload. It's signed, not hidden. Decoding one yourself You can decode the payload in the browser console: const [, payload ] = token . split ( " . " ); console . log ( JSON . parse ( atob ( payload ))); Or, if you just want to paste a token and instantly see the header + payload (without sending it to a server), I built a free decoder that runs entirely in your browser: https://forgly.dev/tools/jwt-decoder Decoding ≠ verifying Reading a JWT is trivial. Trusting it is not — you must verify the signature on your server with the secret/public key before relying on any claim. Decoding just lets you inspect what's there. That's the whole mental model: a JWT is a readable, signed envelope — not a locked box.

2026-05-31 原文 →
AI 资讯

Mobile won the platform war on distribution, not capability

Author here. Wrote this from the vantage point of having shipped Android apps from 2010 to 2019-20. It's a platform-war retrospective along an axis I don't see articulated cleanly very often: update friction and channel control rather than runtime capability. Short version of the argument: The native-only residue (where mobile genuinely wins on capability) is thinner than the narrative claims. UPI in India because the device is the channel. Frame-budget and AR-heavy games. Sustained background GPS. RAW camera. Delivery/on-demand, with a tell (the apps are richer than the web versions because that's where the channel control is, not because the web cannot do it). Electron is the keystone, not a defensive aside. Slack, VS Code, Postman, Bruno, Spotify desktop. If web-versus-native were the deciding axis, the entire web-shell desktop category should have failed. It did not, because the desktop channel is open and the maintainer ships on their own cadence. PWAs are the reverse proof. Apple controls three brakes on the iOS web channel: the WebKit-only rule, the buried Add to Home Screen, and the notification permission. iOS web push did not land until 16.4 in March 2023, years after Android. When the channel is suppressed, the maintainer's update advantage does not save you. The Android hobbyist economy died not from a market outcome but from a channel outcome. Cambridge Analytica 2018 was the public license for a multi-year platform-hardening cycle (target-SDK floor, scoped storage, Play Integrity, foreground-service mandates) that progressively re-priced what kind of software was even shippable. The store does not just control the install button; the channel itself keeps narrowing on the people inside it. It's Part 1 of a two-part series on delivery channels. Posting because I'd rather have the argument tested here than not. submitted by /u/lordVader1138 [link] [留言]

2026-05-31 原文 →
AI 资讯

007 First Light is already discounted for the PS5 and Steam

IO Interactive’s 007 First Light is here, and it’s just as stunning a James Bond mov — err, video game — as we hoped it would be. Pardon the confusion, the title’s engaging tutorial really feels like you’re watching a great Bond movie at times. Whether you’re a longtime Hitman fan who’s been eagerly waiting […]

2026-05-31 原文 →