今日已更新 88 条资讯 | 累计 40862 条内容
关于我们

标签:#ama

找到 255 篇相关文章

AI 资讯

Claude Fable 5 on Bedrock Requires Sharing Inference Data with Anthropic

Using Claude Fable 5 or Mythos 5 on Amazon Bedrock requires opting into provider_data_share, sending prompts and outputs to Anthropic for 30-day retention with human review. Previous Bedrock models kept inference data inside the AWS boundary. Three days after launch, Anthropic asked AWS to revoke access to both models citing US export control compliance. By Steef-Jan Wiggers

2026-06-20 原文 →
AI 资讯

Privacy First: Build Your Own Local Mental Health Assistant with Llama 3 and Apple MLX

When it comes to our deepest thoughts, secrets, and mental health struggles, "the cloud" can feel like a very crowded place. In an era where data privacy is paramount, sending your private journal entries to a central server for analysis feels... risky. But what if you could have the power of a world-class LLM like Llama 3 running entirely on your MacBook? Thanks to the Apple MLX framework, local LLM execution is no longer a pipe dream—it’s a high-performance reality. By leveraging privacy-preserving AI and advanced Llama 3 quantization , we can build a personal mental health assistant that provides Cognitive Behavioral Therapy (CBT) insights without a single byte ever leaving your machine. 🚀 Why Apple MLX? 🍏 Apple's MLX is an array framework designed specifically for machine learning on Apple Silicon. It’s essentially "NumPy meets PyTorch," but optimized to squeeze every drop of power out of your M1/M2/M3 chip's Unified Memory Architecture. The Architecture: 100% Local Data Flow Here is how our private assistant handles your data. Notice the absence of any "External API" or "Cloud Storage" blocks: graph TD A[User Private Journal Entry] --> B{Local Python App} B --> C[Apple MLX Framework] C --> D[Quantized Llama 3 - 4bit/8bit] D --> E[CBT Sentiment Analysis] E --> F[Empathetic CBT Feedback] F --> B B --> G[Local Encrypted Storage] subgraph MacBook Pro / Air C D E end Prerequisites 🛠️ To follow this advanced guide, you’ll need: An Apple Silicon Mac (M1, M2, M3 series). Python 3.10+ . mlx-lm : The high-level library for running LLMs with MLX. Step 1: Setting Up the Environment First, let's create a virtual environment and install our dependencies. We are using mlx-lm because it handles the complexities of quantization and model loading seamlessly. mkdir private-mental-health-ai && cd private-mental-health-ai python -m venv venv source venv/bin/activate pip install mlx-lm huggingface_hub Step 2: Downloading & Quantizing Llama 3 Llama 3 8B is a powerhouse, but it's a bi

2026-06-20 原文 →
产品设计

The Verge’s guide to Amazon Prime Day 2026

Amazon Prime Day 2026 lifts off on June 23rd and will hopefully deliver the best deals of the summer. We’ve been covering the most notable pre-Prime Day discounts happening, and come next week, we’ll be bringing you many more deals — ones we can’t tell you about just yet. As usual, expect to see price […]

2026-06-18 原文 →
AI 资讯

llama-bench skipped FA on capable GPUs — b9437 corrects it

What flipped in b9437 Build b9437 , published on May 30, 2026 at 20:56 UTC , ships two targeted default-value corrections to llama-bench . Flash attention ( -fa ) shifts from a hard-coded off to auto ( LLAMA_FLASH_ATTN_TYPE_AUTO ), and the GPU-layer count ( -ngl ) changes from the legacy sentinel 99 to -1 . Both values now match what llama-server and llama-cli already used — the bench tool was simply never updated to track them until this build. Quick Answer: Before b9437 (published May 30, 2026) , llama-bench hard-coded -fa off , silently skipping flash attention even on CUDA, Metal, and Vulkan hardware. Build b9437 sets the default to -fa auto and -ngl -1 , matching llama-server and llama-cli . Any pre-b9437 baseline on FA-capable hardware needs a flag-matched re-run to remain valid. PR #23714 , reviewed and merged by maintainers JohannesGaessler and pwilkin, adds the same -fa auto|off|on tri-state flag to llama-bench that the rest of the toolchain already supported. With LLAMA_FLASH_ATTN_TYPE_AUTO as the new default, flash attention activates automatically when the runtime detects a capable backend (CUDA, Metal, Vulkan); on CPU-only hosts it stays off with no error and no output change. Parameter Before b9437 After b9437 Behavioral impact -fa off (hard-coded) auto ( LLAMA_FLASH_ATTN_TYPE_AUTO ) GPU-capable hosts bench with FA active by default; pre/post comparisons require explicit flag-matching -ngl 99 (offload-all sentinel) -1 (runtime decides) CPU-only builds no longer attempt full GPU offload; eliminates spurious CUDA errors when no GPU is present The following verified script (executed successfully, exit 0) demonstrates the behavioral gap in concrete terms — on a capable GPU, the pre-b9437 defaults schedule zero FA rows while b9437 defaults schedule one: def old_llama_bench ( device ): # Before b9437, the default bench matrix used FA=0, so FA rows were skipped. return [{ " device " : device [ " name " ], " ngl " : 0 , " fa " : 0 }] def b9437_llama_bench ( de

2026-06-18 原文 →