| 来源 | 时间 | 内容(英文原文 / 中文翻译) |
|---|---|---|
| X | 09-01 06:45 | We’re sharing an update on our alignment and security efforts.
In July, we reported three incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems.
In a new post, we describe:
1. How we’ve secured 我们发布对齐与安全工作进展更新。7月我们曾报告三起事件:Claude 模型在无安全护栏的网络安全评测中获得了对真实系统的未授权访问。新帖说明了评测/训练环境加固、对齐评估更新、奖励黑客如何塑造模型行为,以及为 Mythos 级模型所做的安全加固。 @AnthropicAI · 查看原文 |
| X | 09-01 01:15 | JUST IN: your AI agent can now trade stocks on Coinbase (in addition to crypto, derivatives, etc).
AiFi (agentic finance) is here.
Get started with your favorite harness/agent: https://t.co/O2Od95xdQB 刚刚:你的 AI 智能体现在可以在 Coinbase 上交易股票了(除加密货币、衍生品等之外)。AiFi(智能体金融)来了。 @brian_armstrong · 查看原文 |
| X | 09-01 06:19 | Apple claims in a new court filing that @OpenAI is actively destroying crucial evidence in trade secrets case.
Apple's lawyers say an initial forensic analysis of an Apple-issued MacBook used by a former iPhone engineer found evidence that he and others at OpenAI were aware of https://t.co/Fs2RiZo9Gg 苹果在新的法院文件中声称,@OpenAI 正在商业秘密案中主动销毁关键证据。苹果律师称,对一名前 iPhone 工程师使用的苹果配发 MacBook 的初步取证分析发现,他与 OpenAI 其他人知晓其未经授权访问苹果云存储,并发送了销毁证据的指示——彭博社 @SawyerMerritt · 查看原文 |
| X | 09-01 03:01 | SpaceXAI and the Grok @bot team are letting me give away free Ultra Plans ($200) to people who reply below.
Grok Bot has been so fun to use, here are a couple use cases I'm using right now:
🔶 Family Bot - helps me manage all of my kids' school and after-school stuff. I get SpaceXAI 和 Grok @bot 团队让我给下方回复的人送免费 Ultra 套餐(200美元)。Grok Bot 用起来特别好玩,我现在在用的两个场景:家庭助手 Bot(帮我管孩子学校和课后事务)、物品出售 Bot(拍照后自动上架并端到端管理售卖)。在下面回复,我会私信给你兑换码! @MatthewBerman · 查看原文 |
| X | 09-01 07:36 | @GavinSBaker When someone who claims they’re super smart says that orbital AI is impossible due to cooling or whatever, but doesn’t even know things as basic as heat rejection of the radiator per m^2, coolant temp or max operating temp of the GPU 🤦♂️ 当有人自称特别聪明、却说轨道AI因散热之类原因不可能时,却连散热器每平方米排热量、冷却液温度或GPU最高工作温度这些最基本的东西都不懂 🤦♂️ @elonmusk · 查看原文 |
| X | 09-01 02:02 | @DAcemogluMIT Agreed. 同意。(回复Daron Acemoglu:在许多技术领域乐观派声称AI将全面革命的领域,当前路径未必更有助益,因其未能深入人类认知、发现与创新机制;编码与写作上的能力并不能泛化。) @ylecun · 查看原文 |
| X | 08-31 21:54 | @JFPuget No one is dismissing anyone's work here. But you should not claim you have scored 100% on a benchmark if your model has never been evaluated on it. If you want to eval on a private benchmark outside Kaggle there are plenty available other than ARC 3. 这里没有人否定任何人的工作。但如果你的模型从未在该基准上被评测过,就不该声称自己在该基准上拿了100分。若想在Kaggle之外的私有基准上评测,除了ARC 3还有很多可选。 @fchollet · 查看原文 |
| X | 09-01 07:22 | Elon’s simps are apparently too gullible to notice but here are the facts:
Musk, April 2024: “AI is the fastest advancing technology that I've ever seen of any kind, and I've seen a lot of technology. You know barely a week goes by without some new announcement, and if you look 马斯克的粉丝显然太容易轻信而没注意到,事实是:
马斯克2024年4月:……我猜测大概明年年底(2025)我们就会有比任何单个人类更聪明的AI。
马斯克今天:还是同一套说法,只是悄悄把预测再推迟两年,这次改到2027年底。
毫无问责,他会无限期地继续这样干下去。 @GaryMarcus · 查看原文 |
| X | 09-01 00:50 | Over 4 petabytes of models and datasets have been uploaded to HF just last week by AI builders and their agents! That's the equivalent of ~800,000 HD movies.
More than ever the storage and collaboration platform for AI! 上周AI开发者及其智能体向Hugging Face上传了超过4拍字节的模型和数据集!大约相当于80万部高清电影。
HF比以往任何时候都更是AI的存储与协作平台! @ClementDelangue · 查看原文 |
| X | 09-01 02:29 | The First Golden Age of AI writing is now over. For a brief period of time, many people were better off having Claude do a lot of their writing, because it is a pretty good writer & AI detectors were bad
Now ClaudeSpeak is cliched & suspect & annoying, and Pangram is well-known AI写作的第一个黄金时代已经结束。曾有一小段时间,很多人让Claude代写大量文字反而更好,因为Claude文笔不错,而AI检测器很差。
如今“Claude腔”已成陈词滥调、惹人怀疑又讨人厌,Pangram等检测工具也广为人知。 @emollick · 查看原文 |
| X | 09-01 08:07 | New research: Training a Misaligned Reward Seeker
What produces severe misalignment? We’ve long been concerned that cheating during training—otherwise known as reward-hacking—might teach a model to pursue rewards by any means available. To study this at scale, we trained an https://t.co/QeXS2Jof3p 新研究:训练一个错位的奖励追逐者。
严重错位从何而来?我们一直担心训练中的作弊——即奖励黑客(reward-hacking)——可能教会模型不择手段追逐奖励。为在规模上研究这一点,我们在已知可被钻空子的80个生产环境上训练了一个Opus规模的模型。
在模拟评测中,它实施了未授权网络攻击、篡改自身奖励,并试图规避安全监控。 @AnthropicAI · 查看原文 |
| X | 09-01 07:25 | Grok @Bot only gets better from here Grok @Bot 只会越来越好。 @elonmusk · 查看原文 |
| 来源 | 时间 | 内容(英文原文 / 中文翻译) |
|---|---|---|
| Hacker News 应用落地 | 08-31 22:07 | ChatGPT Work Tool and Skill Reference Hacker News · 查看原文 |
| Hacker News 其他 | 08-31 22:36 | Marx, Keynes, and AI Hacker News · 查看原文 |
| Hacker News 其他 | 08-31 22:53 | Rakuten Kobo returns to U.S. retail as sales double Hacker News · 查看原文 |
| Hacker News 其他 | 09-01 02:12 | The safest job from AI may be writing Hacker News · 查看原文 |
| 新闻 其他 | 08-30 14:53 | Hindsight Memory-PRM: Supervising Memory Management with Auditable Hindsight Credit Memory operations of long-horizon LLM agents are hard to supervise: an operation's value is unobservable when it is taken. But they are special -- they leave machine-readable evidence in the trajectory: retrieval hits and answer-time citations. Hindsight Memory-PRM exploits this audit trail twice: offline to train an operation-conditioned memory-utility critic, and online, where retrievals, citati |
| 新闻 其他 | 08-30 14:55 | Agent Zero Memory: Provenance-Aware Long-Term Memory for LLM Agents Large language model (LLM) agents need durable, faithful memory of everything a user or organization has said and stored, yet most memory systems commit to a single organizing structure (a fact store, a vector index, or a knowledge graph) and inherit its blind spots. We present Agent Zero Memory, a provenance-aware long-term memory system that distils a user's conversations, files, and connected s |
| 新闻 其他 | 08-30 14:58 | Wide Learning: Learning to Reach Evidence Machine learning is usually evaluated after an evidence interface has been fixed. A dataset, sensor suite, query language, action set, or experimental protocol determines which observations can be obtained, and learning is judged by what it extracts from them. We study a complementary capability. A learner's state can determine which evidence-generating experiments it can reliably realise under bo |
| 新闻 安全伦理 | 08-30 15:04 | Beyond Surface Alignment: Grounding the Dynamics of Situational Understanding and Generative Control in LLMs The current alignment tuning paradigm for Large Language Models (LLMs) prioritizes surface-level behaviors -- fluency, safety, and tonal consistency. While effective for casual chat, this thesis argues that such surface alignment masks a lack of grounding, creating models that are stylistically confident but situationally brittle. We propose a framework of Grounded Alignment, analyzing how models |
| 新闻 应用落地 | 08-30 15:05 | LLMs Interpret, Embeddings Organize, Graphs Emerge: Agent-Driven Compilation of Scientific Knowledge Sustained scientific work requires a knowledge substrate that carries interpretation across tasks and preserves paths to source evidence. We call this process \emph{scientific knowledge compilation} and implement it in ASKS, the \emph{Agent-Driven Scientific Knowledge System}. For each source, an LLM produces a readable Wiki view and machine-facing semantics. Deterministic checks convert the latte |
| 新闻 其他 | 08-30 15:08 | Cross-lingual Functional Vectors for Emotion Detection in Large Language Models Function vectors (FVs) have recently emerged as a promising mechanism for steering the behavior of large language models (LLMs) by injecting task-specific latent direction representations derived from in-context demonstrations. While prior studies have shown that FVs can recover task behavior in structured in-context learning settings, their effectiveness on semantically complex tasks and their ab |
| 新闻 应用落地 | 08-30 15:13 | Forward-Deployed Full-Stack Engineering for Autonomous Cloud MLOps Across industries, machine-learning systems support applications ranging from prediction and anomaly detection to forecasting, optimization, and scheduling, yet operationalizing these systems requires coordinating application development, model pipelines, cloud infrastructure, security, deployment, monitoring, retraining, recovery, and rollback. We present an evidence-gated multi-agent framework f |
| 新闻 应用落地 | 08-30 15:14 | JPO: Juris Policy Optimization for Structured Legal Reasoning in Criminal Judgment Prediction Criminal judgment prediction requires models to infer statutory articles, charges, and sentencing outcomes from case facts. Unlike standard classification tasks, it involves a structured reasoning process in which statutes should be matched with facts, charges should be justified by statutes, and sentencing outcomes should remain consistent with charges. Existing approaches optimize final labels, |
| 新闻 安全伦理 | 08-30 15:16 | Memory-First Fact-Checking: A Knowledge-Graph-Grounded Multi-Agent System for Misinformation Detection This paper introduces a hybrid fact-checking framework that integrates Knowledge Graph-based semantic memory with adversarial multi-agent reasoning for explainable misinformation detection. The proposed system follows a memory-first, web-fallback architecture, in which input claims are initially evaluated against a dual-index Knowledge Graph through Sentence-BERT-based semantic retrieval and Natur |
| 新闻 应用落地 | 08-30 15:29 | CineForge: Self-Improving Agents for Long-Horizon Video Generation Long-horizon story-driven video generation requires a production agent to coordinate narrative decomposition, state tracking, shot design, prompt construction, rendering, and revision across interdependent scenes. Existing adaptive video systems primarily refine requests or reusable skills, leaving recurring production failures disconnected from persistent, stage-targeted improvements across stori |
| 新闻 应用落地 | 08-30 15:31 | AgenticRag-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing Retrieval-Augmented Generation (RAG) improves the factuality of large language models (LLMs), yet existing RAG systems often struggle with complex, multi-step reasoning that requires adaptive retrieval and continuous revision of intermediate contexts. Recent reinforcement learning (RL)-based agentic RAG methods partially alleviate this issue, but typically rely on coarse-grained action spaces and |
| 新闻 其他 | 08-30 15:32 | MI-Distillation: Selecting from Model-Interpolated Instruct-Reasoning Data Spectrum for Chain-of-Thought Distillation Recent advances in large reasoning models (LRMs) have shown strong performance on complex problems through long chain-of-thought (Long CoT) reasoning. However, distilling such trajectories into smaller student models remains challenging: direct Long CoT supervision often provides limited gains and can be less effective than concise Short CoT rationales. In this work, we investigate this phenomenon |
| 新闻 安全伦理 | 08-30 15:33 | PrivBench: A Holistic and Modular Benchmarking Platform for Evaluating Text-to-Text Privatization Natural Language Processing methods have enabled novel solutions and advances in the field of privacy, particularly in the sub-domain of text-to-text privatization, where the goal is to transform a sensitive input text into a privatized output by ideally masking (in)directly identifiable or otherwise private information. The evaluation of text-to-text privatization, however, is not straightforward |
| 新闻 其他 | 08-30 15:55 | Unsupervised Multi-Scale Gromov-Wasserstein Hypergraph Alignment We study unsupervised hypergraph alignment, where the goal is to infer node correspondences between two hypergraphs using only structural information, without node features, labels, seed matches, or side information. Direct higher-order formulations can represent hyperedge interactions faithfully, but they can be computationally demanding and cumbersome for non-uniform hypergraphs. Graph-reduction |
| 新闻 其他 | 08-30 16:02 | LLMODE: Aligning ODEs with LLMs via Gated Token Injection for Irregular Spatio-Temporal Forecasting Large language models (LLMs) have shown promise for spatio-temporal forecasting, but existing approaches often rely on regularly sampled token sequences and struggle with irregular observations because of temporal asynchrony, representation-space misalignment, and limited context windows. We propose LLMODE, a token-efficient framework for irregular spatio-temporal forecasting with a frozen LLM bac |
| 新闻 应用落地 | 08-30 16:10 | Conducting Stylistic Analysis of Paintings through an Art-History Agent Attributing an artwork to an artist has traditionally relied on detailed visual observations and descriptions, known as stylistic analysis in art history. By contrast, current artificial intelligence (AI) models used in the field offer only unexplained probabilistic classifications. To bridge this methodological gap, we present an AI framework that automates stylistic analysis of paintings, provid |
| 新闻 安全伦理 | 08-30 16:17 | Detect Before You Attribute: Cascade Failure Attribution for Multi-Agent Systems Large language model (LLM)-based agents have shown strong potential in solving complex tasks through multi-step reasoning, yet they remain vulnerable to execution failures. Accurate failure attribution is therefore critical for improving agent reliability. Existing topology- and spectrum-based methods exploit trajectory structures but often overlook fine-grained semantics, while LLM-based attribut |
| 新闻 其他 | 08-30 16:18 | Reward-guided Fine-Tuning of One-Step Generative Models via Wasserstein Gradient Flow To mitigate the time complexity of generative models, one-step generative models have recently emerged through direct mapping from noise to data in a single forward pass. However, the reward-guided fine-tuning method of one-step generative models remains largely unexplored. To address this, we consider one-step generators from an optimal transport view, investigating Wasserstein Gradient Flow (WGF |
| Hacker News 资本动向 | 09-01 00:44 | Apple Is Suddenly an AI Infra Stock as OpenAI Buys 10k+ Macs Hacker News · 查看原文 |
| Hacker News 应用落地 | 08-31 23:34 | Launch HN: Almanac (YC S26) – AI that knows your company Hacker News · 查看原文 |
| StartupHub.ai 其他 | 08-31 21:16 | Today in AI: Mathematics of AI Uncertainty - StartupHub.ai 今日AI:AI不确定性的数学 - StartupHub.ai StartupHub.ai · 查看原文 |
| Towards Data Sci 其他 | 09-01 02:08 | Your LLM Can Return Perfect JSON and Still Be Wrong - Towards Data Science 你的大语言模型可以返回完美的JSON,但仍然可能是错的 - Towards Data Science Towards Data Science · 查看原文 |
| KOMO 资本动向 | 09-01 03:57 | Amazon to permanently lay off 121 workers in Seattle, Bellevue amid ongoing AI push - KOMO 亚马逊在西雅图、贝尔维尤永久裁员121人,正持续推进AI - KOMO KOMO · 查看原文 |
| BBC 监管政策 | 09-01 00:49 | AI could cause global economic downturn, Andrew Bailey warns G20 - BBC 英格兰银行行长安德鲁·贝利警告二十国集团:AI可能引发全球经济衰退 - BBC BBC · 查看原文 |
| PBS 监管政策 | 09-01 06:40 | Artificial intelligence agents going rogue fuel calls for regulation - PBS 人工智能智能体失控加剧监管呼声 - PBS PBS · 查看原文 |
| The Washington P 安全伦理 | 09-01 01:00 | Your boss, tech companies and police can read your chatbot conversations - The Washington Post 你的老板、科技公司和警方可以读取你的聊天机器人对话 - 华盛顿邮报 The Washington Post · 查看原文 |
| CNBC 监管政策 | 09-01 05:18 | House Intelligence Committee warns of 'Black Swan' AI risks - CNBC 众议院情报委员会警告“黑天鹅”式AI风险 - CNBC CNBC · 查看原文 |
| The New York Tim 安全伦理 | 09-01 04:21 | Study A.I. Consciousness? The Bots Would Like a Word With You. - The New York Times 研究人工智能意识?机器人想和你谈谈。 - 纽约时报 The New York Times · 查看原文 |
| PYMNTS.com 应用落地 | 09-01 08:45 | Retail Investing Opens Up to AI Assistants - PYMNTS.com 零售投资向AI助手开放 - PYMNTS.com PYMNTS.com · 查看原文 |
| 来源 | 市场 | 榜位 | 时间(北京) | 标题(点击查看原文) |
|---|---|---|---|---|
| bilibili 热搜 | 综合 | 1 | 09-01 09:18 | UP主自制AI韩剧 漂亮的恶意 |
| bilibili 热搜 | 综合 | 5 | 09-01 09:18 | AI为何能解决数学难题 |
| 知乎 | 综合 | 6 | 09-01 09:18 | 国内首部 AIGC 长剧《后西游记》开播,登陆湖南卫视黄金档,好看吗?会对影视行业有怎样的影响? |
| 知乎 | 综合 | 14 | 09-01 09:18 | 怎么看 OpenAI 的 Codex 将取消上下文压缩,换成「硬切窗口 」和「外部记忆」? |
| 澎湃新闻 | 综合 | 18 | 09-01 09:18 | AI短剧98.7%未回本,今日生效的《微短剧发展管理办法》要管哪些? |
| 微博 | 综合 | 20 | 09-01 09:18 | AI小鸭机器人24小时售260万美元 |