{
	"updated": "2026-09-05",
	"updated_at": "2026-09-05 13:05 PDT",
	"days": [
		{
			"date": "2026-09-05",
			"items": [
				{
					"rank": 1,
					"title": "RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning",
					"title_zh": "RoboTok：面向人类演示检索与灵巧操作学习的互联网规模数据引擎",
					"title_ja": "RoboTok: 人間のデモ検索と器用な操作学習のためのインターネット規模データエンジン",
					"url": "https://huggingface.co/papers/2609.03199",
					"arxiv_id": "2609.03199",
					"upvotes": 62,
					"first_author": "Howard Qian",
					"org": "Rice RobotPI Lab",
					"summary": "RoboTok is a data engine that takes a query video of a human manipulating an object and retrieves matching demonstrations from web video at internet scale, then trains dexterous robot policies on them. It learns a latent motion space from 3D hand trajectories in actor-centered reference frames, so retrieval transfers across people and scenes. The target is the long tail of tasks robot data collection cannot cover.",
					"summary_zh": "RoboTok 是一个数据引擎：给定一段人类操作物体的查询视频，它能在互联网规模的网络视频中检索出相关的演示片段，用于训练灵巧操作机器人策略。该方法从以操作者为中心坐标系表示的三维手部轨迹中学习潜在运动空间，使检索能够跨越不同人物与场景。其目标是覆盖机器人数据采集难以负担的长尾任务。",
					"summary_ja": "RoboTok は、人が物体を操作するクエリ動画を与えると、インターネット規模のウェブ動画から関連するデモを検索し、器用な操作を行うロボット方策の学習に用いるデータエンジンだ。動作主中心の座標系で表した3次元手指軌跡から潜在的な運動空間を学習するため、人物や場面が変わっても検索が機能する。ロボットのデータ収集では賄いきれない長い裾野のタスクを狙う。"
				},
				{
					"rank": 2,
					"title": "Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction",
					"title_zh": "Scal3R：为可扩展在线三维重建学习高效的多参考相对位姿查询",
					"title_ja": "Scal3R: スケーラブルなオンライン3D再構成のための効率的な多参照相対姿勢クエリの学習",
					"url": "https://huggingface.co/papers/2609.04201",
					"arxiv_id": "2609.04201",
					"upvotes": 45,
					"first_author": "Chin-Yang Lin",
					"org": "National Yang Ming Chiao Tung University",
					"summary": "Online 3D reconstruction degrades on long videos, and the authors trace the failure to regressing every pose against a fixed first-frame anchor, which pushes the model far outside its training distribution. Per-frame depth stays intact while the global pose head collapses, so Scal3R reframes the task as querying poses relative to multiple references. The change keeps reconstruction stable over much longer sequences.",
					"summary_zh": "在线三维重建在长视频上会明显退化，作者将失败归因于把所有位姿都相对于固定的首帧锚点回归，这会让模型远远偏离训练分布。研究发现逐帧深度依然可靠，崩溃的只是全局位姿头，因此 Scal3R 把任务重新表述为相对多个参考帧的位姿查询。这一改动让重建在更长的序列上保持稳定。",
					"summary_ja": "オンライン3D再構成は長い動画で精度が崩れるが、著者らはその原因を、すべての姿勢を固定した先頭フレームを基準に回帰させるため学習分布から大きく外れる点に求めた。フレームごとの深度は保たれ、崩壊するのは大域的な姿勢推定部分だけであることから、Scal3R は複数の参照フレームに対する相対姿勢のクエリとして問題を定式化し直す。これにより、はるかに長い系列でも再構成が安定する。"
				},
				{
					"rank": 3,
					"title": "Editable Visual Design",
					"title_zh": "可编辑视觉设计",
					"title_ja": "編集可能なビジュアルデザイン",
					"url": "https://huggingface.co/papers/2609.04034",
					"arxiv_id": "2609.04034",
					"upvotes": 40,
					"first_author": "Junyan Ye",
					"org": "Tencent Hunyuan",
					"summary": "Diffusion image models produce flattened bitmaps with error-prone text and no way to edit a layer afterwards, while code-based generation gives clean layers but weak aesthetic judgment. The paper proposes a coding-agent paradigm in which a vision-language model acts as the creative brain and emits code, keeping layout control and editable layers. The 12-author work comes from Tencent Hunyuan.",
					"summary_zh": "扩散图像模型生成的是压平的位图，文字容易出错，事后也无法按图层修改；而基于代码的生成虽然图层清晰，却缺乏整体审美判断。论文提出一种以编码智能体为核心的范式，让视觉语言模型充当创意大脑并输出代码，从而同时保留布局控制与可编辑图层。这项由 12 位作者完成的工作来自腾讯混元。",
					"summary_ja": "拡散モデルによる画像生成は文字が崩れやすい平坦なビットマップを出力し、後からレイヤー単位で編集できない。一方、コードによる生成はレイヤーが整理される反面、全体の美的判断が弱い。本論文は、視覚言語モデルを創造の頭脳としてコードを生成させるコーディングエージェント型の枠組みを提案し、レイアウト制御と編集可能なレイヤーを両立させる。著者12名による腾訊混元（Tencent Hunyuan）の成果。"
				},
				{
					"rank": 4,
					"title": "The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation",
					"title_zh": "缺失的时间链接：面向剧本驱动音视频生成的时间上下文路由",
					"title_ja": "欠けていた時間的リンク: 脚本駆動の音声・映像生成のための時間文脈ルーティング",
					"url": "https://huggingface.co/papers/2609.02367",
					"arxiv_id": "2609.02367",
					"upvotes": 33,
					"first_author": "Yichen Liu",
					"summary": "Joint audio-video generators have gotten good at lip and scene synchronization but still offer little control over when a shot cuts or a line of dialogue lands, which breaks script-driven production. The authors show that timing written into a structured prompt survives only in the text representation and never reaches the generator's temporal axis. Their temporal context routing carries it through explicitly.",
					"summary_zh": "音视频联合生成模型在口型与画面同步上已经相当成熟，但对镜头在何时切换、台词在何时出现仍缺乏控制，这使其难以用于剧本驱动的内容创作。作者指出，结构化提示词中写明的时间信息只存在于文本表示里，从未真正传递到生成模型的时间轴上。他们提出的时间上下文路由则把这些信息显式地传导过去。",
					"summary_ja": "音声と映像を同時に生成するモデルは同期の品質を高めてきたが、いつショットが切り替わり、いつ台詞が発せられるかの制御は依然として乏しく、脚本に沿った制作には使いにくい。著者らは、構造化プロンプトに書かれた時間指定がテキスト表現の中に留まり、生成側の時間軸まで届いていないことを示す。提案する時間文脈ルーティングは、その情報を明示的に受け渡す。"
				},
				{
					"rank": 5,
					"title": "Last Translation Benchmark",
					"title_zh": "最后的翻译基准",
					"title_ja": "ラスト翻訳ベンチマーク",
					"url": "https://huggingface.co/papers/2609.04173",
					"arxiv_id": "2609.04173",
					"upvotes": 25,
					"first_author": "Vilém Zouhar",
					"summary": "A 244-author collaboration argues that standard machine translation benchmarks are approaching saturation while automatic metrics remain unreliable, open to reward hacking and unactionable, and even gold human evaluation lacks reproducibility and scale. The paper offers a benchmark and evaluation method built to test the limits of current models and surface specific failure cases rather than emit one number.",
					"summary_zh": "这项由 244 位作者共同完成的工作指出，主流机器翻译基准正逼近饱和，自动指标既不可靠、易被奖励攻击，也无法指导改进，而被视为金标准的人工评测同样缺乏可复现性与规模。论文提出一套基准与评测方法，目标是逼出当前模型的能力上限并暴露具体失败案例，而不是给出一个分数。",
					"summary_ja": "244名の共同研究は、標準的な機械翻訳ベンチマークが飽和に近づく一方、自動評価指標は信頼性を欠き報酬ハッキングに弱く改善の指針にもならず、金標準とされる人手評価さえ再現性と規模に乏しいと論じる。本論文は、単一のスコアを出すのではなく、現行モデルの限界を突き、具体的な失敗事例を可視化するためのベンチマークと評価手法を示す。"
				}
			],
			"updated_at": "2026-09-05 13:05 PDT"
		},
		{
			"date": "2026-09-04",
			"items": [
				{
					"rank": 1,
					"title": "Compile by Training: Turning Natural-Language Specifications into Local Neural Functions",
					"url": "https://huggingface.co/papers/2609.04199",
					"arxiv_id": "2609.04199",
					"upvotes": 230,
					"first_author": "Yuntian Deng",
					"org": "University of Waterloo",
					"title_zh": "以训练完成编译：将自然语言规范转化为本地神经函数",
					"title_ja": "訓練によるコンパイル: 自然言語仕様をローカルな神経関数へ",
					"summary": "Researchers at the University of Waterloo propose compile by training, which turns a natural-language specification into a reusable neural function. Teacher models generate task-specific examples at compile time to train a small adapter for a compact interpreter, so the result runs without the teachers and can be stored, versioned and composed like ordinary software. It leads the day's papers with 230 votes.",
					"summary_zh": "滑铁卢大学的研究者提出以训练完成编译的方法，把自然语言规范转化为可复用的神经函数。编译阶段由教师模型生成任务示例，用于训练紧凑解释器上的小型适配器，所得函数此后无需教师模型即可运行，并可像普通软件一样存储、版本管理与组合。该论文以 230 票位居当日榜首。",
					"summary_ja": "ウォータールー大学の研究者らは、自然言語の仕様を再利用可能な神経関数に変換する「訓練によるコンパイル」を提案した。コンパイル時に教師モデルがタスク固有の例を生成し、小型インタプリタ向けのアダプタを学習させるため、完成した関数は教師なしで動作し、通常のソフトウェアと同様に保存・版管理・合成ができる。230票で当日首位となった。"
				},
				{
					"rank": 2,
					"title": "Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments",
					"url": "https://huggingface.co/papers/2609.04148",
					"arxiv_id": "2609.04148",
					"upvotes": 209,
					"first_author": "Jie Wu",
					"org": "Qwen",
					"title_zh": "Terminal-Universe：把智能体轨迹转化为可扩展的终端环境",
					"title_ja": "Terminal-Universe: エージェントの軌跡をスケール可能な端末環境へ",
					"summary": "A Qwen team reconstructs executable terminal environments from the tool-execution history recorded in existing agent trajectories. The authors argue that post-training needs environments rather than trajectories, since each environment can be re-queried into many verifiable tasks and returns execution feedback, while a trajectory is a single frozen demonstration. The paper drew 209 votes.",
					"summary_zh": "Qwen 团队从已有智能体轨迹中记录的工具执行历史，重建可运行的终端环境。作者认为后训练真正需要的是环境而非轨迹：每个环境都能被反复提问以生成大量可验证任务并给出执行反馈，而一条轨迹只是一次固定的示范。该论文获得 209 票。",
					"summary_ja": "Qwenのチームは、既存のエージェント軌跡に残るツール実行履歴から、実行可能な端末環境を再構築した。著者らは事後学習に必要なのは軌跡ではなく環境だと論じる。環境は何度も問い直して検証可能なタスクを多数生み出し実行フィードバックを返すが、軌跡は固定された一度きりの実演にすぎないためだ。209票を集めた。"
				},
				{
					"rank": 3,
					"title": "Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training",
					"url": "https://huggingface.co/papers/2608.26730",
					"arxiv_id": "2608.26730",
					"upvotes": 139,
					"first_author": "Tingyun Li",
					"title_zh": "知道何时不该复用：自主大模型后训练中的条件化经验迁移",
					"title_ja": "再利用しない判断: 自律的なLLM事後学習における条件付き経験転移",
					"summary": "This paper asks which past update evidence still applies once later training has changed the parent model. In autonomous post-training systems that propose updates, train candidates and select from evaluation feedback, an update's effect depends on its parent, data and training stage, so treating earlier success as context-free permission wastes compute. It collected 139 votes.",
					"summary_zh": "该论文追问：当后续训练已改变父模型后，过去的更新证据还有多少仍然适用。在自动提出更新、训练候选并依据评测反馈择优的后训练系统中，一次更新的效果取决于其父模型、数据与训练阶段，把早先的成功视为无条件的通行证会浪费算力。该论文获得 139 票。",
					"summary_ja": "本論文は、その後の学習で親モデルが変化した時点で、過去の更新の証拠がどこまで有効なのかを問う。更新を提案し候補を学習させ評価フィードバックで選ぶ自律的な事後学習システムでは、更新の効果は親モデル・データ・学習段階に依存するため、以前の成功を無条件の許可と見なせば計算資源を浪費する。139票を集めた。"
				},
				{
					"rank": 4,
					"title": "Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning",
					"url": "https://huggingface.co/papers/2609.03430",
					"arxiv_id": "2609.03430",
					"upvotes": 108,
					"first_author": "Heng Wang",
					"org": "Salesforce AI Research",
					"title_zh": "随机注意力：重新审视高效推理中的 KV 缓存淘汰",
					"title_ja": "ランダム・アテンション: 効率的推論のためのKVキャッシュ削減を問い直す",
					"summary": "Salesforce AI Research reports that the scoring signal behind KV cache eviction contributes almost nothing. Their Random Attention keeps the prompt and evicts uniformly at random within each attention head, computing no score at all, and across four models and six reasoning tasks it matches the strongest prior evictor. The paper drew 108 votes.",
					"summary_zh": "Salesforce AI Research 发现，KV 缓存淘汰所依赖的打分信号几乎没有贡献。其提出的随机注意力保留提示词，在每个注意力头内均匀随机淘汰，完全不计算分数，在四个模型与六项推理任务上仍可比肩此前最强的淘汰方法。该论文获得 108 票。",
					"summary_ja": "Salesforce AI Researchは、KVキャッシュ削減の土台となるスコア信号がほとんど寄与していないと報告した。提案手法ランダム・アテンションはプロンプトを残し、各アテンションヘッド内で一様にランダムに破棄するだけでスコアを一切計算しないが、4モデル・6種の推論タスクで従来最良の手法に並んだ。108票を集めた。"
				},
				{
					"rank": 5,
					"title": "It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning",
					"url": "https://huggingface.co/papers/2609.00638",
					"arxiv_id": "2609.00638",
					"upvotes": 69,
					"first_author": "Runpeng Dai",
					"org": "Apple",
					"title_zh": "匹配需要双方：用强化学习协同演进的生成式检索器",
					"title_ja": "マッチングは双方向: 強化学習で共進化する生成的リトリーバー",
					"summary": "Apple researchers present CoGR, which trains language models to build retrieval representations on both the query side and the item side rather than using generation only to expand queries. Retrieval is the first stage of search and advertising systems, selecting candidates for downstream ranking and auction. The paper drew 69 votes.",
					"summary_zh": "苹果的研究者提出 CoGR，训练语言模型同时在查询侧与物品侧构建检索表示，而非仅用生成来扩写查询。检索是搜索与广告系统的第一环，负责从海量物品中筛出候选集，供下游排序与竞价使用。该论文获得 69 票。",
					"summary_ja": "Appleの研究者らはCoGRを提案し、生成をクエリ拡張だけに使うのではなく、クエリ側とアイテム側の双方で検索表現を構築するよう言語モデルを学習させた。検索は検索広告システムの最初の段階であり、後段のランキングとオークションに渡す候補集合を選び出す。69票を集めた。"
				}
			],
			"updated_at": "2026-09-04 13:05 PDT"
		},
		{
			"date": "2026-09-03",
			"items": [
				{
					"rank": 1,
					"title": "HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?",
					"url": "https://huggingface.co/papers/2609.01437",
					"arxiv_id": "2609.01437",
					"upvotes": 178,
					"first_author": "Yuhao Wu",
					"org": "ByteDance Seed",
					"title_zh": "HarnessDev：大模型能否创建并演进自己的智能体框架？",
					"title_ja": "HarnessDev: LLMは自らのエージェント・ハーネスを構築し進化させられるか",
					"summary": "ByteDance Seed introduces HarnessDev, a benchmark that evaluates models on building and evolving the agent harness itself rather than on task outputs. The unit of evaluation becomes runnable infrastructure, split across two stages. The premise is that changing the harness while holding weights fixed can substantially alter task performance, an ability current evaluations leave underexplored.",
					"summary_zh": "字节跳动 Seed 团队提出 HarnessDev 基准，评测对象不是任务输出，而是模型自行构建与演进智能体框架（harness）的能力。评测单位由此变为可运行的基础设施，分为两个阶段。其前提是：在权重不变的情况下改动 harness 就能显著改变任务表现，而现有评测对这一能力探索不足。",
					"summary_ja": "ByteDance Seed は、タスクの出力ではなくエージェント・ハーネス自体を構築・進化させる能力を測るベンチマーク HarnessDev を提案した。評価の単位は実行可能なインフラそのもので、2段階で構成される。重みを固定したままハーネスを変えるだけでタスク性能が大きく変わるという前提に立ち、既存評価が手薄な領域を突く。"
				},
				{
					"rank": 2,
					"title": "Language Models Can Control Their Own Attention",
					"url": "https://huggingface.co/papers/2609.02737",
					"arxiv_id": "2609.02737",
					"upvotes": 50,
					"first_author": "Namgyu Ho",
					"org": "KAIST AI",
					"title_zh": "语言模型可以自行控制注意力",
					"title_ja": "言語モデルは自らの注意を制御できる",
					"summary": "KAIST AI proposes an intrinsic route to sparse attention: rather than pre-selecting context tokens with lightweight external proxy scores, the model itself signals which parts of the context matter. The target is the cost of global attention layers reading the full KV cache at every step, which proxy scoring still pays at O(N) per step.",
					"summary_zh": "KAIST AI 提出一种内生的稀疏注意力思路：不再用轻量的外部代理分数预先挑选上下文 token，而是让模型自己指示哪些上下文重要。针对的是全局注意力层每生成一个 token 都要读取完整 KV 缓存的开销，而代理打分方案每步仍需 O(N) 的代价。",
					"summary_ja": "KAIST AI は、外部の軽量な代理スコアで文脈トークンを事前選別するのではなく、モデル自身がどの部分が重要かを示す内在的なスパース注意の手法を提案する。狙いは、グローバル注意層が毎ステップKVキャッシュ全体を読む負荷であり、代理スコア方式でも1ステップあたりO(N)の計算が残る点にある。"
				},
				{
					"rank": 3,
					"title": "H3-World: Turning Language Understanding into World Control",
					"url": "https://huggingface.co/papers/2609.01560",
					"arxiv_id": "2609.01560",
					"upvotes": 47,
					"first_author": "Danze Chen",
					"title_zh": "H3-World：把语言理解转化为世界控制",
					"title_ja": "H3-World: 言語理解を世界制御へと変える",
					"summary": "H3-World turns the 33B MiniMax-H3 video generator into an interactive world model without adding dedicated action modules. Each action is represented as a structured language instruction, sharpening the generator's existing zero-shot control of character behavior and camera motion into precise, temporally grounded control.",
					"summary_zh": "H3-World 将 330 亿参数的 MiniMax-H3 视频生成模型改造为可交互的世界模型，且无需引入专门的动作模块。方法是把每个动作表示为结构化的语言指令，从而把该模型原有的零样本角色行为与镜头运动控制，细化为精确且在时间上对齐的控制。",
					"summary_ja": "H3-World は、330億パラメータの動画生成モデル MiniMax-H3 を、専用の行動モジュールを追加せずに対話的なワールドモデルへと転換する。各行動を構造化された言語指示として表現することで、同モデルが元から持つキャラクター挙動やカメラ運動のゼロショット制御を、時間的に整合した精密な制御へと洗練させる。"
				},
				{
					"rank": 4,
					"title": "ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training",
					"url": "https://huggingface.co/papers/2609.00188",
					"arxiv_id": "2609.00188",
					"upvotes": 43,
					"first_author": "Xionghao Wu",
					"org": "Joy Future Academy",
					"title_zh": "ZimaBlue：通过可扩展的视频预训练演化通用世界动作模型",
					"title_ja": "ZimaBlue: スケーラブルな動画事前学習による汎用ワールドアクションモデルの進化",
					"summary": "ZimaBlue learns generalizable world action models for robotic manipulation from egocentric video instead of action-labeled robot trajectories. It addresses a scaling gap: robust generalization needs broad physical experience, yet labeled trajectories are expensive and narrow. Action-free video supplies contact dynamics, tool use and long-horizon behavior across diverse settings.",
					"summary_zh": "ZimaBlue 从第一人称视角视频中学习可泛化的机器人操作世界动作模型，而非依赖带动作标注的机器人轨迹。它针对的是规模化难题：稳健泛化需要广泛的物理经验，但标注轨迹成本高且多样性有限。无动作标注的视频则能提供接触动力学、工具使用与长时程行为等多场景经验。",
					"summary_ja": "ZimaBlue は、行動ラベル付きのロボット軌跡ではなく一人称視点の動画から、汎化するロボット操作用ワールドアクションモデルを学習する。頑健な汎化には広範な身体的経験が要るが、ラベル付き軌跡は高コストで多様性に乏しいという規模の壁に取り組む。行動ラベルのない動画が、接触の力学や道具の使用、長期的な振る舞いを多様な環境から供給する。"
				},
				{
					"rank": 5,
					"title": "From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix",
					"url": "https://huggingface.co/papers/2609.01572",
					"arxiv_id": "2609.01572",
					"upvotes": 33,
					"first_author": "Olga Tsymboi",
					"org": "T-Tech",
					"title_zh": "从生产流量到后训练：构建覆盖企业请求分布的自托管大模型",
					"title_ja": "本番トラフィックから事後学習へ: 企業の要求構成を網羅する自己ホスト型LLMの構築",
					"summary": "T-Tech reports consolidating traffic from more than 200 internal applications onto a single self-hosted model. Quality gaps were found through production error analysis along three axes - instruction following, function calling and internal task distribution - and tracked with offline benchmarks stratified to production traffic. The work targets the GPU fragmentation that data-residency rules create.",
					"summary_zh": "T-Tech 介绍了如何把 200 多个内部应用的流量整合到单一自托管模型上。团队通过生产环境的错误分析，从指令遵循、函数调用与内部任务分布三个维度定位质量差距，并用按生产流量分层的离线基准进行跟踪。其目标是缓解数据驻留要求所造成的 GPU 资源碎片化。",
					"summary_ja": "T-Tech は、200を超える社内アプリケーションのトラフィックを単一の自己ホスト型モデルへ統合した取り組みを報告する。品質の差は本番環境のエラー分析から、指示追従・関数呼び出し・社内タスク分布の3軸で特定し、本番トラフィックに層化したオフラインベンチマークで追跡した。データ所在地規制が生むGPUの分断が主な課題である。"
				}
			],
			"updated_at": "2026-09-03 13:06 PDT"
		},
		{
			"date": "2026-09-02",
			"items": [
				{
					"rank": 1,
					"title": "StudentSim: Training LLM-based Student Simulators",
					"url": "https://huggingface.co/papers/2609.01591",
					"upvotes": 220,
					"arxiv_id": "2609.01591",
					"first_author": "Ke Yang",
					"org": "Microsoft Research",
					"title_zh": "StudentSim: 训练基于大语言模型的学生模拟器",
					"title_ja": "StudentSim: LLM ベースの学習者シミュレータを訓練する",
					"summary": "Microsoft Research introduced StudentSim, a training framework for LLM-based simulators that stand in for real learners when evaluating AI tutors. It targets a gap between state-tracking models, which fit student behavior but handle explanations and corrections poorly, and LLM role-play, which follows guidance fluently without matching the imitated student's competence. It is the day's most upvoted paper.",
					"summary_zh": "微软研究院提出 StudentSim，用于训练基于大语言模型的学生模拟器，在评估 AI 辅导系统时替代真实学习者。该工作针对两类方法之间的空白: 状态追踪模型能拟合学生行为，却难以处理讲解与纠正; 大模型角色扮演能流畅遵循指导，却无法匹配被模仿学生的真实水平。这是当日票数最高的论文。",
					"summary_ja": "Microsoft Research は、AI チューターの評価において実在の学習者の代わりを務める、LLM ベースの学習者シミュレータを訓練する枠組み StudentSim を発表した。狙うのは既存手法の隙間だ。状態追跡モデルは行動を再現できても説明や訂正の処理が苦手で、LLM のロールプレイは指導に流暢に従う一方、模倣対象の学力水準を再現できない。本日最多の投票を集めた論文である。"
				},
				{
					"rank": 2,
					"title": "DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution",
					"url": "https://huggingface.co/papers/2608.31106",
					"upvotes": 93,
					"arxiv_id": "2608.31106",
					"first_author": "Jiashu Zhu",
					"org": "AMAP-ML",
					"title_zh": "DreamX-Creator: 让 2K 分辨率的原生音视频生成走向普及",
					"title_ja": "DreamX-Creator: 2K 解像度のネイティブ音声・映像同時生成を身近に",
					"summary": "AMAP-ML presented DreamX-Creator 1.0, a compact system built on a 7B generator that denoises audio and video jointly instead of adding sound in a separate stage. Conditioned on a first frame and a text prompt, the two streams run independently through the first half of the network and are coupled in the latter half by gated cross-modal attention.",
					"summary_zh": "AMAP-ML 发布 DreamX-Creator 1.0，一套以 70 亿参数生成器为核心的紧凑系统，音频与视频联合去噪，而非在独立阶段后补声音。模型以首帧和文本提示为条件，两路数据流在网络前半段各自独立处理，后半段通过门控跨模态注意力耦合在一起。",
					"summary_ja": "AMAP-ML は、7B の生成器を中核とするコンパクトな音声・映像同時生成システム DreamX-Creator 1.0 を発表した。音声を別段階で後付けするのではなく、両者を同時にノイズ除去する。先頭フレームとテキストプロンプトを条件とし、2 つのストリームはネットワーク前半では独立に処理され、後半でゲート付きクロスモーダル注意により結合される。"
				},
				{
					"rank": 3,
					"title": "GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling",
					"url": "https://huggingface.co/papers/2608.29335",
					"upvotes": 60,
					"arxiv_id": "2608.29335",
					"first_author": "Guangting Zheng",
					"org": "ByteDance Seed",
					"title_zh": "GenFirst: 生成先于重建, 实现稳定的端到端隐空间生成建模",
					"title_ja": "GenFirst: 再構成より先に生成を——安定した端から端までの潜在生成モデリング",
					"summary": "ByteDance Seed revisited end-to-end latent generative modeling, where the autoencoder and the generator are trained together rather than fitting a generator on a frozen, reconstruction-optimized latent space. Analyzing how different objectives shape that space, the authors target the latent collapse and the generation-reconstruction conflict that have made joint training unstable.",
					"summary_zh": "字节跳动 Seed 团队重新审视端到端的隐空间生成建模: 自编码器与生成模型联合训练，而不是在冻结的、为重建而优化的隐空间上再训练生成模型。作者分析了不同训练目标如何塑造隐空间，针对导致联合训练不稳定的隐空间坍缩与生成-重建冲突提出解法。",
					"summary_ja": "ByteDance Seed は、端から端までの潜在生成モデリングを再検討した。再構成向けに最適化して凍結した潜在空間の上で生成モデルを学習するのではなく、オートエンコーダと生成器を同時に訓練する枠組みだ。著者らは各目的関数が潜在空間をどう形作るかを分析し、同時学習を不安定にしてきた潜在崩壊と生成・再構成の対立に取り組む。"
				},
				{
					"rank": 4,
					"title": "Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving",
					"url": "https://huggingface.co/papers/2609.00111",
					"upvotes": 56,
					"arxiv_id": "2609.00111",
					"first_author": "Xin Zhou",
					"org": "Qwen",
					"title_zh": "Qwen-Drive-1.0: 迈向自动驾驶视觉语言基础模型的第一步",
					"title_ja": "Qwen-Drive-1.0: 自動運転向け視覚言語基盤モデルへの第一歩",
					"summary": "Qwen released Qwen-Drive-1.0, a vision-language foundation model for autonomous driving that keeps the pretrained VLM architecture and folds 3D perception, visual question answering and motion planning into one framework. An external bird's-eye-view head performs 3D detection, occupancy prediction and map segmentation, probing what 3D structure the shared representations expose.",
					"summary_zh": "Qwen 发布 Qwen-Drive-1.0，一个面向自动驾驶的视觉语言基础模型，保留预训练视觉语言模型的架构，并将三维感知、视觉问答与运动规划整合进统一框架。外接的鸟瞰图头部负责三维目标检测、语义占据预测与地图分割，用以探查共享表征中究竟包含多少三维结构信息。",
					"summary_ja": "Qwen は、自動運転向けの視覚言語基盤モデル Qwen-Drive-1.0 を公開した。事前学習済み VLM のアーキテクチャを保ちつつ、3D 認識・視覚質問応答・動作計画を単一の枠組みに統合する。外付けの鳥瞰図ヘッドが 3D 物体検出、占有予測、地図セグメンテーションを担い、共有表現からどの程度の 3D 構造が取り出せるかを探る役割も果たす。"
				},
				{
					"rank": 5,
					"title": "SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers",
					"url": "https://huggingface.co/papers/2609.01343",
					"upvotes": 53,
					"arxiv_id": "2609.01343",
					"first_author": "Shaowen Wang",
					"org": "ByteDance Seed",
					"title_zh": "SMELT: 算力对齐条件下 MoE 循环 Transformer 的扩展律",
					"title_ja": "SMELT: 計算量を揃えた MoE ループ型 Transformer のスケーリング則",
					"summary": "ByteDance Seed measured looped Transformers under matched compute, holding per-token FLOPs, non-embedding parameters and KV cache constant so that extra depth is not confused with extra FLOPs. The resulting SMELT recipe loops the middle half of a sparse mixture-of-experts model's layers twice, and is scaled across four sizes up to 54B non-embedding parameters.",
					"summary_zh": "字节跳动 Seed 团队在算力对齐的条件下评估循环 Transformer: 固定每 token 的 FLOPs、非嵌入参数量与 KV 缓存，避免把额外深度带来的收益与额外算力混为一谈。由此得到的 SMELT 方案将稀疏专家混合模型中间一半的层循环两次，并在最高 540 亿非嵌入参数的四个规模上验证。",
					"summary_ja": "ByteDance Seed は、計算量を揃えた条件でループ型 Transformer を評価した。トークンあたりの FLOPs、非埋め込みパラメータ数、KV キャッシュを一致させ、深さの増加による効果と計算量の増加を混同しないようにしている。得られた SMELT のレシピは、疎な MoE モデルの中間半分の層を 2 回ループさせるもので、非埋め込み 540 億パラメータまでの 4 規模で検証された。"
				}
			],
			"updated_at": "2026-09-02 01:04 PDT"
		},
		{
			"date": "2026-09-01",
			"items": [
				{
					"rank": 1,
					"title": "Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement",
					"url": "https://huggingface.co/papers/2608.31046",
					"upvotes": 90,
					"arxiv_id": "2608.31046",
					"first_author": "Yi Ding",
					"org": "Purdue University",
					"title_zh": "同策略蒸馏真的在蒸馏吗？从含噪教师到自我提升",
					"title_ja": "オンポリシー蒸留は本当に蒸留しているのか: ノイズを含む教師から自己改善へ",
					"summary": "Purdue researchers measured the supervision a teacher supplies during on-policy distillation, where the teacher scores trajectories that are off-policy for it, and found substantial noise that becomes more prevalent as the teacher grows larger. The student policy converges regardless, which the authors read as pointing to self-improvement rather than transfer from the teacher.",
					"summary_zh": "普渡大学的研究者量化了同策略蒸馏中教师提供的监督信号：教师需要给对它而言属于离策略的学生轨迹打分，结果发现其中噪声显著，且教师规模越大噪声越普遍。学生策略却对这种噪声并不敏感、照常收敛，作者据此认为提升更多来自自我改进而非来自教师的知识迁移。",
					"summary_ja": "パデュー大学の研究者らは、オンポリシー蒸留において教師が与える監督信号を定量的に分析した。教師は自身にとってはオフポリシーである生徒の軌跡を採点するため、その信号には無視できないノイズが含まれ、教師の規模が大きいほど頻度が増す。それでも生徒の方策は収束しており、著者らは性能向上の源が教師からの転移ではなく自己改善にあると読み解いている。"
				},
				{
					"rank": 2,
					"title": "Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling",
					"url": "https://huggingface.co/papers/2608.30821",
					"upvotes": 65,
					"arxiv_id": "2608.30821",
					"first_author": "Minghan Qin",
					"org": "ByteDance Seed",
					"title_zh": "Lucida：面向可组合真实到仿真场景建模的解析、生成与摆放",
					"title_ja": "Lucida: 構成可能な実世界からシミュレーションへのシーンモデリングのための解析・生成・配置",
					"summary": "ByteDance Seed presents Lucida, a pipeline that turns a capture of a real indoor room into complete, individually editable object assets arranged as they were observed. It targets the weak point of existing parse-generate-place systems, each stage of which assumes accurate instance geometry, unoccluded views and assets that match the scene. The output is aimed at robot simulation and embodied AI.",
					"summary_zh": "字节跳动 Seed 团队提出 Lucida，可将真实室内场景的采集数据还原为完整且可单独编辑的物体资产，并按原有布局摆放。该工作针对现有「解析-生成-摆放」流程的薄弱环节：每一步都预设了准确的实例几何、无遮挡视角以及与场景吻合的资产，而杂乱的实拍很难满足。成果面向机器人仿真与具身智能。",
					"summary_ja": "ByteDance Seed は、実際の室内空間の撮影データを、観測どおりに配置された編集可能な個別オブジェクト資産へと復元するパイプライン Lucida を提案した。既存の「解析・生成・配置」型手法は各段階で正確なインスタンス形状、遮蔽のない視点、シーンに合致した資産を前提とするが、雑然とした実撮影ではそれが成り立たない点を突く。出力はロボットのシミュレーションや身体性 AI での利用を想定する。"
				},
				{
					"rank": 3,
					"title": "J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data",
					"url": "https://huggingface.co/papers/2608.26582",
					"upvotes": 37,
					"arxiv_id": "2608.26582",
					"first_author": "Gyouk Chu",
					"org": "KAIST AI",
					"title_zh": "J-Zero：从零数据出发的出题者-解题者-评判者统一协同进化",
					"title_ja": "J-Zero: ゼロデータから始める出題者・解答者・審判の統合的共進化",
					"summary": "KAIST AI trains a challenger, a solver and a judge together from no seed data, so a model can keep improving in domains where an answer cannot be checked automatically. The challenger raises task difficulty as the solver adapts to it, while the judge co-adapts alongside them to supply the reward signal that unverifiable tasks lack.",
					"summary_zh": "KAIST AI 让出题者、解题者与评判者在零起始数据的条件下共同训练，使模型也能在答案无法自动校验的领域持续提升。出题者不断加大任务难度、解题者随之调整，评判者则与二者协同演化，为这类不可验证任务补上缺失的奖励信号。",
					"summary_ja": "KAIST AI は、出題者・解答者・審判を初期データなしで同時に学習させ、答えを自動検証できない領域でもモデルが改善し続けられるようにした。出題者が課題の難度を上げ、解答者がそれに適応する一方、審判も両者と並んで共進化し、検証不能な課題に欠けている報酬信号を供給する。"
				},
				{
					"rank": 4,
					"title": "Normalized Low-Rank Adaptation",
					"url": "https://huggingface.co/papers/2608.31036",
					"upvotes": 37,
					"arxiv_id": "2608.31036",
					"first_author": "Jiale Kang",
					"title_zh": "归一化低秩适配",
					"title_ja": "正規化された低ランク適応",
					"summary": "NoRA normalizes the down-projection matrices during LoRA training, on the observation that a zero-initialized up-projection leaves early optimization almost entirely to the down-projection. The authors report that applying the same normalization only at initialization already improves standard LoRA, making it a drop-in change to a widely used fine-tuning method.",
					"summary_zh": "NoRA 在 LoRA 训练过程中对下投影矩阵做归一化，其依据是：上投影初始化为零，因此早期优化动态几乎完全由下投影主导。作者称，即便只在初始化阶段施加同样的归一化，也能改善标准 LoRA，相当于为这一常用微调方法提供了可直接替换的改动。",
					"summary_ja": "NoRA は、LoRA の学習中にダウン射影行列を正規化する手法である。アップ射影がゼロで初期化されるため、初期の最適化はほぼダウン射影に委ねられるという観察に基づく。著者らは、同じ正規化を初期化時にのみ適用しても標準的な LoRA が改善すると報告しており、広く使われる微調整手法にそのまま差し込める変更となる。"
				},
				{
					"rank": 5,
					"title": "StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments",
					"url": "https://huggingface.co/papers/2608.24804",
					"upvotes": 35,
					"arxiv_id": "2608.24804",
					"first_author": "Esakkivel Esakkiraja",
					"org": "ServiceNow-AI",
					"title_zh": "StarHarness：以分层搜索为企业环境演化智能体框架",
					"title_ja": "StarHarness: 層化探索によるエンタープライズ環境向けハーネスの進化",
					"summary": "ServiceNow's StarHarness evolves the scaffolding around an agent, including task framing, tool interfaces, skills, MCP-backed providers, subagent structure and agent-loop configuration, while model weights stay fixed. It stratifies tasks by baseline failure behavior and separates proposer-visible tasks from those held back for selection and evaluation, tested on IT operations and finance benchmarks.",
					"summary_zh": "ServiceNow 的 StarHarness 在模型权重保持不变的前提下，演化智能体外围的框架，涵盖提示与任务表述、工具接口、技能、由 MCP 提供的服务、子智能体结构以及智能体循环配置。它按基线失败情况对任务分层，并把提议者可见的搜索任务与用于筛选和留出评测的任务分开，在 IT 运维与金融基准上进行了验证。",
					"summary_ja": "ServiceNow の StarHarness は、モデルの重みを固定したまま、エージェントを取り巻く足場、すなわちプロンプトや課題の提示、ツールのインタフェース、スキル、MCP 経由のプロバイダ、サブエージェント構成、エージェントループの設定を進化させる。課題をベースラインでの失敗の仕方で層化し、提案側から見える探索用の課題と、選択および評価用に取り置く課題を分離したうえで、IT 運用や金融のベンチマークで検証している。"
				}
			],
			"updated_at": "2026-09-01 13:05 PDT"
		},
		{
			"date": "2026-08-31",
			"items": [
				{
					"rank": 1,
					"title": "LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering",
					"title_zh": "LoopArena：将模型作为循环工程运行时控制器的基准测试",
					"title_ja": "LoopArena: ループエンジニアリングの実行時コントローラーとしてのモデルを評価する",
					"url": "https://huggingface.co/papers/2608.28281",
					"arxiv_id": "2608.28281",
					"upvotes": 82,
					"first_author": "Yi Wang",
					"org": "AMAP-ML",
					"summary": "A seven-author team at AMAP-ML introduces LoopArena, a benchmark that scores a model as the runtime controller of a coding-agent loop rather than as the coder. The authors note that the outcome of one end-to-end run cannot separate the loop's guidance from the agent's own ability, so the benchmark isolates decisions such as when to verify, where to spend budget and when to stop.",
					"summary_zh": "AMAP-ML 的七人团队提出 LoopArena，这一基准考察模型作为编码智能体循环的运行时控制器，而非作为写代码的一方。作者指出，单次端到端运行的结果无法区分成败源于循环的调度还是智能体自身能力，因此基准把何时验证、预算投向何处、何时停止等决策单独拆出评估。",
					"summary_ja": "AMAP-ML の 7 名のチームは、コーディングエージェントのループを制御する実行時コントローラーとしてモデルを採点するベンチマーク LoopArena を提案した。著者らは、1 回のエンドツーエンド実行の結果ではループの誘導とエージェント自身の能力を切り分けられないと指摘し、いつ検証するか、予算をどこに使うか、いつ停止するかといった判断を切り出して評価する。"
				},
				{
					"rank": 2,
					"title": "DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents",
					"title_zh": "DART-SD：面向多轮工具调用智能体自蒸馏的菱形拓扑感知检索与调优",
					"title_ja": "DART-SD: 多ターンのツール呼び出しエージェントの自己蒸留に向けたダイヤモンド構造を考慮した検索とチューニング",
					"url": "https://huggingface.co/papers/2608.18524",
					"arxiv_id": "2608.18524",
					"upvotes": 60,
					"first_author": "Hangrui Xu",
					"org": "ByteDance",
					"summary": "ByteDance researchers propose DART-SD, a self-distillation recipe for multi-turn tool-calling agents. They argue that imitating full-length trajectories collapses the combinatorial lattice formed by order-independent sub-goals, indiscriminately penalizing valid alternative orderings and degrading policy diversity. DART-SD instead pairs topology-aware retrieval with tuning to keep that structure.",
					"summary_zh": "字节跳动的研究者提出 DART-SD，一种面向多轮工具调用智能体的自蒸馏方法。他们认为，模仿完整长度的轨迹会压垮由次序无关子目标构成的组合格结构，对同样有效的其他执行顺序一并施加惩罚，从而削弱策略多样性。DART-SD 转而以拓扑感知的检索配合微调来保留这一结构。",
					"summary_ja": "バイトダンスの研究者らは、多ターンのツール呼び出しエージェント向けの自己蒸留手法 DART-SD を提案した。全長の軌跡をそのまま模倣すると、順序に依存しないサブゴールが作る組合せ格子が潰れ、妥当な別順序まで一律に罰せられて方策の多様性が損なわれると論じる。DART-SD は構造を考慮した検索とチューニングを組み合わせてこれを保つ。"
				},
				{
					"rank": 3,
					"title": "Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090",
					"title_zh": "Puro-2B：小实验室在 RTX 5090 上以 5090 美元训练的 Qwen2-1.5B",
					"title_ja": "Puro-2B: 小規模ラボが RTX 5090 上で 5090 ドル以内に学習した Qwen2-1.5B",
					"url": "https://huggingface.co/papers/2608.27370",
					"arxiv_id": "2608.27370",
					"upvotes": 26,
					"first_author": "Kairong Luo",
					"org": "PACMAN Group, Tsinghua University",
					"summary": "An 11-author team from Tsinghua University's PACMAN Group publishes an open pretraining recipe and the Puro-2B models it produces, trained on an RTX 5090 for under 5,090 dollars. The report notes that training Llama-3.2-3B costs over 1.5 million dollars and reproducing SmolLM3-3B over 700,000, and argues a cost-efficient, hardware-accessible recipe has been the missing piece for academic and open-source groups.",
					"summary_zh": "清华大学 PACMAN 组的十一人团队公开了一套预训练配方以及由此训练出的 Puro-2B 系列模型，全部在一块 RTX 5090 上完成，成本低于 5090 美元。报告指出，训练 Llama-3.2-3B 需超过 150 万美元，复现 SmolLM3-3B 也要 70 万美元以上，而学术界与开源社区一直缺少一份低成本、硬件可及的配方。",
					"summary_ja": "清華大学 PACMAN グループの 11 名のチームが、オープンな事前学習レシピと、それにより RTX 5090 上で 5,090 ドル未満で学習した Puro-2B 群を公開した。報告は Llama-3.2-3B の学習に 150 万ドル超、SmolLM3-3B の再現に 70 万ドル超がかかると指摘し、低コストで入手しやすいハードウェア向けのレシピが学術・オープンソース側に欠けていたと論じる。"
				},
				{
					"rank": 4,
					"title": "Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models",
					"title_zh": "意图先行：为视觉-语言-动作模型蒸馏行为意图",
					"title_ja": "意図を持って行動する：視覚・言語・行動モデルのための行動意図の蒸留",
					"url": "https://huggingface.co/papers/2608.23478",
					"arxiv_id": "2608.23478",
					"upvotes": 24,
					"first_author": "Sangoh Lee",
					"org": "Pohang University of Science and Technology",
					"summary": "Three researchers at POSTECH propose Intention Distillation for vision-language-action models, whose action decoders are still trained largely by behavior cloning. That supervision records which motor command was demonstrated but leaves the local objective implicit, and the authors argue future-based signals such as frames or trajectories capture particular realizations rather than the shared aim of the behavior.",
					"summary_zh": "韩国浦项科技大学的三位研究者为视觉-语言-动作模型提出意图蒸馏，此类模型的动作解码器目前仍主要依靠行为克隆训练。这种监督只记录示范中执行了哪条运动指令，而把该行为在指令下所服务的局部目标留作隐含；作者认为基于未来帧或轨迹的信号捕捉的是某一次具体实现，而非行为共有的语义目标。",
					"summary_ja": "韓国・浦項工科大学の 3 名の研究者は、視覚・言語・行動モデル向けに意図蒸留を提案した。これらのモデルの行動デコーダーは今なお主に行動クローニングで学習され、どの運動指令が示範されたかは教えても、その行動が指示のもとで果たす局所的な目的は暗黙のままである。将来のフレームや軌跡に基づく信号は個別の実現例を捉えるにすぎないと著者らは論じる。"
				},
				{
					"rank": 5,
					"title": "LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation",
					"title_zh": "LayerRecall：面向视频生成长时程一致性的状态条件记忆路由器",
					"title_ja": "LayerRecall: 動画生成の長期的一貫性のための状態条件付きメモリルーター",
					"url": "https://huggingface.co/papers/2608.28460",
					"arxiv_id": "2608.28460",
					"upvotes": 24,
					"first_author": "Yixuan Ding",
					"org": "Zhejiang University",
					"summary": "A Zhejiang University team introduces LayerRecall, a memory router for autoregressive video diffusion, which generates long clips chunk by chunk from a bounded recent context and evicts the historical cues needed when a subject or scene reappears. Their analysis finds video DiT layers differ in their preference for current, recent and distant context, so the router decides both what to retrieve and where to use it.",
					"summary_zh": "浙江大学团队提出 LayerRecall，一个用于自回归视频扩散的记忆路由器。此类模型以有限的近期上下文逐块生成长视频，会丢弃主体或场景再次出现时所需的历史线索。团队的分析发现，视频 DiT 的不同层对当前、近期与久远上下文的偏好各异，因此路由器同时决定检索什么以及在哪一层使用。",
					"summary_ja": "浙江大学のチームは、自己回帰的な動画拡散のためのメモリルーター LayerRecall を提案した。この方式は限られた直近の文脈からチャンク単位で長尺動画を生成するため、被写体や場面が再登場する際に必要な過去の手がかりを捨ててしまう。解析の結果、動画 DiT の層ごとに現在・近傍・遠方の文脈への選好が異なると分かり、ルーターは何を取り出すかとどこで使うかを併せて決める。"
				}
			],
			"updated_at": "2026-08-31 13:06 PDT"
		},
		{
			"date": "2026-08-29",
			"items": [
				{
					"rank": 1,
					"title": "Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization",
					"title_zh": "Zero-WAM：基于人类视频的上下文世界-动作建模，实现开放式任务泛化",
					"title_ja": "Zero-WAM: 人間の動画からの文脈内ワールド・アクションモデリングによるオープンエンドなタスク汎化",
					"url": "https://huggingface.co/papers/2608.26103",
					"arxiv_id": "2608.26103",
					"upvotes": 17,
					"first_author": "Jiaming Zhou",
					"org": "Robbyant Research",
					"summary": "Zero-WAM carries in-context learning over from language models to robot manipulation, using a human video rather than a sentence as the specification of a new task. The authors argue video is the natural specification for manipulation because it supplies visual cues that language leaves out. The policy then executes tasks never seen in training without any parameter update.",
					"summary_zh": "Zero-WAM 把大语言模型中的上下文学习迁移到机器人操作上，用一段人类视频而非一句指令来指定新任务。作者认为视频才是操作任务的天然表述方式，因为它提供了语言无法承载的视觉线索。策略据此执行训练中从未见过的任务，且无需更新任何参数。",
					"summary_ja": "Zero-WAM は、大規模言語モデルの文脈内学習をロボット操作に持ち込み、新しいタスクの指定に文ではなく人間の動画を用いる。著者らは、言語が取りこぼす視覚的手がかりを与える動画こそ操作タスクの自然な指定手段だと論じる。方策はこれをもとに、学習時に一度も見ていないタスクをパラメータ更新なしで実行する。"
				},
				{
					"rank": 2,
					"title": "WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution",
					"title_zh": "WikiSkill：将智能体经验编译为持久知识以驱动技能演化",
					"title_ja": "WikiSkill: エージェントの経験を永続的な知識へ編纂しスキルを進化させる",
					"url": "https://huggingface.co/papers/2608.27454",
					"arxiv_id": "2608.27454",
					"upvotes": 11,
					"first_author": "Liyan Tang",
					"org": "Google",
					"summary": "Google proposes WikiSkill, a framework that evolves an agent's skill library alongside a persistent knowledge base rather than in isolation. It targets a specific gap in automatic skill discovery: the insights that guide how a skill develops stay scattered through optimization histories and are rarely reused across iterations. WikiSkill keeps raw execution experience and accumulated knowledge in separate layers.",
					"summary_zh": "谷歌提出 WikiSkill，让智能体的技能库不再孤立演化，而是与一个持久化的知识库（wiki）协同演进。该工作针对自动技能发现中的一处具体缺口：指导技能形成的经验判断散落在优化历史中，难以在多轮迭代间系统复用。WikiSkill 将原始执行经验与沉淀下来的知识分层存放。",
					"summary_ja": "Google は、エージェントのスキル群を単独で更新するのではなく、永続的な知識ベース（wiki）と共進化させるフレームワーク WikiSkill を提案した。自動的なスキル発見における具体的な欠落、すなわちスキルの育て方を導く知見が最適化履歴の中に散在し、反復をまたいで再利用されにくい点を狙う。WikiSkill は生の実行経験と蓄積された知識を層として分離する。"
				},
				{
					"rank": 3,
					"title": "Procedura: Agentic 3D Modeling with Procedural Control",
					"title_zh": "Procedura：具备程序化控制的智能体三维建模",
					"title_ja": "Procedura: 手続き的な制御によるエージェント型 3D モデリング",
					"url": "https://huggingface.co/papers/2608.26238",
					"arxiv_id": "2608.26238",
					"upvotes": 10,
					"first_author": "Youtian Lin",
					"org": "Nanjing University",
					"summary": "Procedura treats a 3D shape as code, scaling an LLM's coding ability to write an object as a procedural assembly whose named parts are joined by typed, machine-checkable mates. The team from Nanjing University aims at three weaknesses of dense meshes from native 3D generators: soft edges where an object should be sharp, no part decomposition, and no parameter a user can edit.",
					"summary_zh": "Procedura 把三维形状当作代码来生成，借助并放大大语言模型的编程能力，将物体写成一段程序化装配：各具名部件之间以带类型、可机器校验的配合关系相连。南京大学团队针对原生三维生成器输出稠密网格的三处短板：本该锐利的边缘变得圆钝、缺乏部件分解、以及不提供任何可供用户修改的参数。",
					"summary_ja": "Procedura は 3D 形状をコードとして扱い、大規模言語モデルのコーディング能力を活かして、名前付きの部品どうしを型付きで機械検証可能な結合関係でつないだ手続き的アセンブリとして物体を記述する。南京大学のチームが狙うのは、ネイティブ 3D 生成器が出す密なメッシュの三つの弱点、すなわち鋭くあるべき箇所の甘さ、部品分解の欠如、そして編集できるパラメータが一切ないことである。"
				},
				{
					"rank": 4,
					"title": "CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes",
					"title_zh": "CritICL：从小语言模型的失败模式出发，实现推理时的由弱到强泛化",
					"title_ja": "CritICL: 小規模言語モデルの失敗モードを用いた推論時のウィーク・トゥ・ストロング汎化",
					"url": "https://huggingface.co/papers/2608.27455",
					"arxiv_id": "2608.27455",
					"upvotes": 8,
					"first_author": "Yufan Wu",
					"summary": "CritICL improves reasoning at inference time without the repeated generation or external verification that inference-time scaling usually requires. Its premise is that failure modes are structured and recur across model scales within one family, so the mistakes of a small model become guidance for a larger one. The authors present failures as a source of signal rather than output to discard.",
					"summary_zh": "CritICL 在推理阶段提升模型的推理能力，却不依赖推理时扩展通常需要的反复采样或外部验证器。其核心假设是：失败模式具有结构性，且在同一模型家族的不同规模之间反复出现，因此小模型犯的错可以转化为对大模型的指引。作者把失败视为可用的信号来源，而非应当丢弃的劣质输出。",
					"summary_ja": "CritICL は、推論時スケーリングが通常必要とする繰り返し生成や外部検証を用いずに、推論時点で推論性能を高める。前提となるのは、失敗の仕方には構造があり、同じモデルファミリー内では規模をまたいで繰り返し現れるという観察であり、小さなモデルの誤りがより大きなモデルへの手がかりになる。著者らは失敗を捨てるべき出力ではなく信号源として位置づける。"
				},
				{
					"rank": 5,
					"title": "Magpie: Real-Time World Renderer for Interactive Games",
					"title_zh": "Magpie：面向交互式游戏的实时世界渲染器",
					"title_ja": "Magpie: インタラクティブゲームのためのリアルタイム世界レンダラー",
					"url": "https://huggingface.co/papers/2608.27168",
					"arxiv_id": "2608.27168",
					"upvotes": 7,
					"first_author": "Xiaoyu Zhan",
					"summary": "Magpie is a real-time generative world renderer built for games rather than linear media. The authors note that video foundation models now reshape film and video work, but a game additionally needs stable, reproducible rules, object states and interaction outcomes on top of continuous imagery. The target is the cost of asset production, which stretches prototype development cycles.",
					"summary_zh": "Magpie 是一个面向游戏而非线性影像的实时生成式世界渲染器。作者指出，视频基础模型正在改变影视与视频制作，但游戏除了连续可信的画面之外，还需要稳定且可复现的规则、物体状态与交互结果。该工作瞄准的是资产制作的高昂成本，这一成本正拉长原型开发周期。",
					"summary_ja": "Magpie は、線形メディアではなくゲームに向けて作られたリアルタイム生成世界レンダラーである。著者らは、動画基盤モデルが映像制作を変えつつある一方で、ゲームには連続した映像に加えて、安定して再現可能なルール、物体の状態、相互作用の結果が必要だと指摘する。狙いは、プロトタイプの開発期間を引き延ばしているアセット制作コストである。"
				}
			],
			"updated_at": "2026-08-29 13:04 PDT"
		},
		{
			"date": "2026-08-28",
			"items": [
				{
					"rank": 1,
					"title": "Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models",
					"title_zh": "以智能体游戏开发作为可验证轨迹数据引擎，用于扩展世界模型",
					"title_ja": "世界モデルのスケーリングに向けた検証可能な軌跡データエンジンとしてのエージェント型ゲーム開発",
					"url": "https://huggingface.co/papers/2608.25518",
					"arxiv_id": "2608.25518",
					"upvotes": 117,
					"first_author": "Pengfei Zhou",
					"org": "National University of Singapore",
					"summary": "The authors argue that scaling world models on more crawled video and more compute is inefficient without grounded reward signals, and propose a recursive data engine built on agentic game development instead. Because games are executable, the trajectories they produce can be scored the way compilers score code agents, rather than through fuzzy proxies such as CLIP scores.",
					"summary_zh": "作者认为，仅靠抓取更多视频、投入更多算力来扩展世界模型效率低下，缺少的是有据可依的奖励信号，并提出改用建立在智能体游戏开发之上的递归式数据引擎。由于游戏可执行，其产生的轨迹能像编译器为代码智能体打分那样获得高质量奖励，而不必依赖 CLIP 分数这类模糊代理指标。",
					"summary_ja": "著者らは、より多くの収集動画と計算資源で世界モデルを拡張する戦略は、根拠のある報酬信号を欠く限り非効率だと論じ、代わりにエージェント型のゲーム開発に基づく再帰的なデータエンジンを提案する。ゲームは実行可能であるため、生成される軌跡はコードエージェントに対する コンパイラのように採点でき、CLIP スコアのような曖昧な代理指標に頼らずに済む。"
				},
				{
					"rank": 2,
					"title": "UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City",
					"title_zh": "UrbanGround：在真实尺度城市中从局部感知走向空间自主行动",
					"title_ja": "UrbanGround: 実スケール都市における局所的知覚から空間的行為主体性へ",
					"url": "https://huggingface.co/papers/2608.27456",
					"arxiv_id": "2608.27456",
					"upvotes": 68,
					"first_author": "Tianjie Ju",
					"org": "Shanghai Jiao Tong University",
					"summary": "UrbanGround tests whether multimodal LLM agents can turn street-level perception into reliable action once they begin to move through a city. The sandbox is a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data, supporting closed-loop first-person interaction and an interactive map for navigation.",
					"summary_zh": "UrbanGround 用于检验多模态大模型智能体在城市中开始移动之后，能否把街景层面的局部感知转化为可靠的行动。该沙盒依据覆盖全境的三维地理空间数据，构建了受物理约束的香港复制体，支持第一人称视角的闭环交互，并提供用于导航的交互式地图。",
					"summary_ja": "UrbanGround は、マルチモーダル LLM エージェントが都市を移動し始めた後も、street view から得た局所的な情報を信頼できる行動に変換できるかを検証する。全域の 3 次元地理空間データから構築した物理制約付きの香港レプリカで、一人称視点の閉ループ操作と、経路探索用の対話的な地図を備える。"
				},
				{
					"rank": 3,
					"title": "TTPO: Test-Time Policy Optimization",
					"title_zh": "TTPO：测试时策略优化",
					"title_ja": "TTPO: テスト時ポリシー最適化",
					"url": "https://huggingface.co/papers/2608.27448",
					"arxiv_id": "2608.27448",
					"upvotes": 62,
					"first_author": "Aozhe Wang",
					"summary": "Post-training methods such as reinforcement learning and on-policy self-distillation depend on ground-truth labels, which rules out test-time training, and majority-vote pseudo-labels are fragile because one wrong vote corrupts the teacher. The authors find the failure is asymmetric: rollouts that disagree with the pseudo-label are usually wrong whether or not the vote itself is, and TTPO builds on that.",
					"summary_zh": "强化学习、同策略自蒸馏等后训练方法依赖真值标签，因而无法用于测试时训练；改用多数投票伪标签又很脆弱，一次投错就会污染教师信号并误导每个 token。作者发现这种失效是不对称的：与伪标签不一致的 rollout 通常本身就是错的，无论投票是否正确，TTPO 正是建立在这一不对称性之上。",
					"summary_ja": "強化学習やオンポリシー自己蒸留といった事後学習は正解ラベルに依存するため、テスト時学習には使えない。多数決の擬似ラベルで代用する手も、一票の誤りが教師信号を汚染し全トークンを誤導するため脆い。著者らはこの失敗が非対称であること、すなわち擬似ラベルと食い違う rollout は投票の正誤にかかわらず概ね誤りである点に着目し、TTPO を構成した。"
				},
				{
					"rank": 4,
					"title": "Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher",
					"title_zh": "Self-OPD：无需教师模型的流匹配模型同策略蒸馏",
					"title_ja": "Self-OPD: 教師モデルを用いないフローマッチングモデルのオンポリシー蒸留",
					"url": "https://huggingface.co/papers/2608.26872",
					"arxiv_id": "2608.26872",
					"upvotes": 55,
					"first_author": "Shiyi Zhang",
					"summary": "On-policy distillation gives dense supervision but normally requires a specialized teacher trained for every new objective, and the gap between teacher and student distributions compounds errors along the generation trajectory. Self-OPD is a teacher-free framework that keeps the dense signal for flow matching models while removing both costs.",
					"summary_zh": "同策略蒸馏能提供稠密的监督信号，但通常要为每个新目标单独训练专用教师模型，成本高昂，而且教师与学生分布之间的差异会沿生成轨迹不断累积误差。Self-OPD 提出一种无需教师的框架，在流匹配模型上保留稠密信号的同时消除这两项开销。",
					"summary_ja": "オンポリシー蒸留は密な教師信号を与えられる一方、新しい目的ごとに専用の教師モデルを訓練する必要があり計算コストが高く、教師と生徒の分布のずれが生成軌跡に沿って誤差を累積させる。Self-OPD は教師を必要としない枠組みで、フローマッチングモデルにおいて密な信号を保ちながらこの二つの問題を取り除く。"
				},
				{
					"rank": 5,
					"title": "What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents",
					"title_zh": "什么才是好的智能体数据？以 ACE 视角审视面向大模型智能体的数据生成",
					"title_ja": "優れたエージェント用データとは何か: LLM エージェント向けデータ生成を ACE の視点で捉える",
					"url": "https://huggingface.co/papers/2608.27260",
					"arxiv_id": "2608.27260",
					"upvotes": 55,
					"first_author": "Xingshan Zeng",
					"summary": "Work on generating interaction data for LLM agents is organized by domain and evaluated inconsistently, which obscures the mechanisms the methods share and blurs candidate construction into verification and selection. This paper proposes a two-level framework for the field, treating consistency among environments, tasks, interactions and success signals as what generation must preserve.",
					"summary_zh": "面向大模型智能体的交互数据生成研究按领域各自组织、评测口径不一，既掩盖了各方法共通的生成机制，也把候选数据构造与验证、筛选混为一谈。本文为该领域提出一个两层框架，把环境、任务、交互与成功信号之间的一致性视作数据生成必须保持的核心。",
					"summary_ja": "LLM エージェント向けの相互作用データ生成の研究は領域ごとに整理され評価もばらばらで、手法に共通する生成メカニズムが見えにくく、候補の構築と検証・選別が混同されている。本論文はこの分野に二層の枠組みを提示し、環境・タスク・相互作用・成功信号の間の一貫性こそ生成が保つべきものだと位置づける。"
				}
			],
			"updated_at": "2026-08-28 13:05 PDT"
		},
		{
			"date": "2026-08-27",
			"items": [
				{
					"rank": 1,
					"title": "VGI-BENCH: Probing Visual Intelligence in Video Generation Models",
					"title_zh": "VGI-BENCH：探测视频生成模型中的视觉智能",
					"title_ja": "VGI-BENCH：動画生成モデルの視覚的知能を測る",
					"url": "https://huggingface.co/papers/2608.19583",
					"summary": "A 22-author team led from the University of Illinois at Urbana-Champaign introduced VGI-bench, a benchmark for the zero-shot visual reasoning that video generation models can show through their generated frames. It holds 27 tasks and 810 instances under a two-level taxonomy of task domains and skill tags. Tasks demand a valid evolving process, not just a plausible final state.",
					"summary_zh": "一支由伊利诺伊大学厄巴纳-香槟分校牵头的 22 人团队提出 VGI-bench，用于评测视频生成模型通过生成帧所展现的零样本视觉推理能力。基准包含 27 类任务、810 个实例，按任务领域与技能标签构成两级分类体系。任务要求模型给出合理的演化过程，而不只是看似合理的最终画面。",
					"summary_ja": "イリノイ大学アーバナ・シャンペーン校を中心とする 22 名の著者チームは、動画生成モデルが生成フレームを通じて示すゼロショットの視覚推論を測るベンチマーク VGI-bench を発表した。タスク領域とスキルタグの 2 層分類のもとに 27 タスク、810 インスタンスを収める。最終状態がもっともらしいだけでは足りず、妥当な過程の推移が求められる。",
					"upvotes": 139,
					"arxiv_id": "2608.19583",
					"first_author": "Xuan He",
					"org": "University of Illinois at Urbana-Champaign"
				},
				{
					"rank": 2,
					"title": "JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution",
					"title_zh": "JIT-Agent：以即时的 harness 演化扩展 harness 智能",
					"title_ja": "JIT-Agent：ジャストインタイムのハーネス進化でハーネス知能をスケールさせる",
					"url": "https://huggingface.co/papers/2608.25593",
					"summary": "Researchers at the National University of Singapore presented JIT-Agent, a model trained to synthesize an agent harness on the fly for any off-the-shelf agentic LLM. It formalizes the harness - memory, planning, action protocol, tool and skill orchestration - as a composable, machine-generatable artifact. The premise is that harness design, still manual, can dominate the contribution of the underlying model.",
					"summary_zh": "新加坡国立大学的研究者提出 JIT-Agent，这是一个经过训练、能为任意现成智能体大模型即时合成 harness 的模型。论文把 harness（记忆管理、规划策略、动作协议与工具及技能编排）形式化为可组合、可由机器生成的产物。其前提是：harness 设计目前仍靠手工且与任务绑定，但它对最终效果的影响可以超过底层模型本身。",
					"summary_ja": "シンガポール国立大学の研究者らは、既製のエージェント LLM 向けにハーネスをその場で合成するよう訓練したモデル JIT-Agent を発表した。記憶管理、計画戦略、行動プロトコル、ツールとスキルの編成から成るハーネスを、組み合わせ可能で機械生成できる成果物として定式化する。ハーネス設計は依然として手作業かつタスク固有だが、その寄与は基盤モデル自体を上回りうるという前提に立つ。",
					"upvotes": 47,
					"arxiv_id": "2608.25593",
					"first_author": "Guibin Zhang",
					"org": "National University of Singapore"
				},
				{
					"rank": 3,
					"title": "CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild",
					"title_zh": "CyberFactory：用真实世界样本扩展网络安全能力",
					"title_ja": "CyberFactory：実世界の事例でサイバーセキュリティ能力をスケールする",
					"url": "https://huggingface.co/papers/2608.23181",
					"summary": "A team at IQuest released CyberFactory, a unified open-source framework for building cybersecurity capability into large language models. The paper argues open efforts trail closed models here because frontier open-weight releases ship no reproducible security training recipe, while existing open work covers isolated tasks and lacks scalable agentic data.",
					"summary_zh": "IQuest 团队发布 CyberFactory，一个用于为大语言模型构建网络安全能力的统一开源框架。论文认为开源阵营在这一方向落后于闭源模型：前沿开放权重模型未提供可复现的安全训练方案，而现有开源工作只覆盖孤立任务，缺少可规模化的智能体数据。",
					"summary_ja": "IQuest のチームは、大規模言語モデルにサイバーセキュリティ能力を持たせるための統一的なオープンソース基盤 CyberFactory を公開した。論文は、フロンティアのオープンウェイト・モデルが再現可能なセキュリティ訓練手順を示さず、既存のオープンな取り組みも個別タスクにとどまり大規模なエージェント・データを欠くため、この領域でクローズドなモデルに後れを取っていると論じる。",
					"upvotes": 30,
					"arxiv_id": "2608.23181",
					"first_author": "Jian Yang",
					"org": "IQuest"
				},
				{
					"rank": 4,
					"title": "D^3-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation",
					"title_zh": "D^3-MOPD：面向高效多教师蒸馏的自适应动态领域调度",
					"title_ja": "D^3-MOPD：効率的な多教師蒸留のための適応的な動的ドメイン・スケジューリング",
					"url": "https://huggingface.co/papers/2608.24987",
					"summary": "A ten-author team proposed D^3-MOPD, which adapts the per-domain data mixture during multi-teacher on-policy distillation instead of fixing it before training. Their motivation is that domains converge at very different rates, so a fixed mixture spends compute on domains that plateaued early while undertraining the slower ones.",
					"summary_zh": "一支十人团队提出 D^3-MOPD：在多教师同策略蒸馏过程中动态调整各领域的数据配比，而不是在训练开始前就将其固定。出发点在于不同领域的收敛速度差异很大，固定配比会把算力浪费在早早收敛的领域，同时让收敛慢的领域训练不足。",
					"summary_ja": "10 名の著者チームは、多教師オンポリシー蒸留において領域ごとのデータ配合を訓練前に固定せず、学習中に適応させる D^3-MOPD を提案した。領域ごとに収束の速さが大きく異なるため、固定配合では早期に頭打ちになった領域に計算資源を費やし、収束の遅い領域は訓練不足になるという問題に対処する。",
					"upvotes": 22,
					"arxiv_id": "2608.24987",
					"first_author": "Zechen Sun"
				},
				{
					"rank": 5,
					"title": "Agent-G^2: Gaussian Guidance for Agentic Reinforcement Learning",
					"title_zh": "Agent-G^2：面向智能体强化学习的高斯引导",
					"title_ja": "Agent-G^2：エージェント強化学習のためのガウス的ガイダンス",
					"url": "https://huggingface.co/papers/2608.23318",
					"summary": "Zhejiang University researchers proposed Agent-G^2 for hint-based reinforcement learning, in which a prefix of an expert trajectory is retained before each rollout so the policy explores from a state closer to success. Existing methods treat that guidance depth as one deterministic number; the authors find useful guidance instead occupies a band of depths.",
					"summary_zh": "浙江大学的研究者提出 Agent-G^2，面向基于提示的强化学习：每次 rollout 前保留一段专家轨迹前缀，让策略从更接近成功的状态开始探索。现有方法把这一「引导深度」当作单一确定值，而作者发现真正有用的引导分布在一个深度区间内。",
					"summary_ja": "浙江大学の研究者らは、ヒント型強化学習に向けた Agent-G^2 を提案した。各ロールアウトの前に熟練軌跡の先頭部分を残し、成功に近い状態から方策を探索させる手法である。既存手法はこのガイダンス深度を単一の決定的な値として扱うが、著者らは有用なガイダンスが一定の深度帯に広がることを見いだした。",
					"upvotes": 20,
					"arxiv_id": "2608.23318",
					"first_author": "Zixuan Wang",
					"org": "Zhejiang University"
				}
			],
			"updated_at": "2026-08-27 13:05 PDT"
		},
		{
			"date": "2026-08-26",
			"items": [
				{
					"rank": 1,
					"title": "GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture",
					"title_zh": "GigaBrain-0.7：用三系统架构将具身基础模型扩展至涌现能力",
					"title_ja": "GigaBrain-0.7: 三システム構成によるエンボディッド基盤モデルのスケーリングと能力の創発",
					"url": "https://huggingface.co/papers/2608.15875",
					"summary": "GigaAI presents GigaBrain-0.7, an embodied foundation model organized around a three-system architecture. The paper asks whether current vision-language-action systems still gain from better architectural design and from far larger, more heterogeneous training data, and reports improved generalization across different robot embodiments and tasks.",
					"summary_zh": "GigaAI 发布具身基础模型 GigaBrain-0.7，采用三系统架构组织。论文追问当前的视觉-语言-动作系统能否从更优的架构设计以及规模更大、更异构的训练数据中继续获益，并报告该模型在不同机器人形态与任务上的泛化能力有所提升。",
					"summary_ja": "GigaAI は、三つのシステムからなる構成を軸としたエンボディッド基盤モデル GigaBrain-0.7 を発表した。論文は、現在の視覚・言語・行動（VLA）システムがより良いアーキテクチャ設計と、はるかに大規模で多様な学習データからなお恩恵を受けられるのかを問い、異なるロボットの身体や課題をまたぐ汎化の改善を報告している。",
					"upvotes": 89,
					"arxiv_id": "2608.15875",
					"first_author": "GigaBrain Team",
					"org": "GigaAI"
				},
				{
					"rank": 2,
					"title": "AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces",
					"title_zh": "AutoSaddler：从智能体执行轨迹中持久更新，自动优化运行框架",
					"title_ja": "AutoSaddler: エージェント実行トレースからの永続的な更新によるハーネス自動最適化",
					"url": "https://huggingface.co/papers/2608.23041",
					"summary": "Microsoft proposes AutoSaddler, which recasts agent harness design as an offline learning problem rather than manual search over prompts, tool configurations and control logic. It reads failure signals from execution traces in mini-batches and iteratively updates the harness. The target is long-horizon tasks, where small local failures compound into overall failure.",
					"summary_zh": "微软提出 AutoSaddler，把智能体运行框架（harness）的设计从对提示词、工具配置与控制逻辑的人工搜索，改写为一个离线学习问题。该方法以小批量方式从执行轨迹中读取失败信号，迭代更新 harness。目标场景是长程任务，其中局部的小失误会累积成整体失败。",
					"summary_ja": "マイクロソフトは AutoSaddler を提案した。プロンプト、ツール構成、制御ロジックを人手で探索していたエージェント・ハーネスの設計を、オフライン学習の問題として定式化する。実行トレースから得た失敗信号をミニバッチ単位で読み取り、ハーネスを反復的に更新する。狙いは、局所的な小さな失敗が積み重なって全体の失敗になる長期タスクである。",
					"upvotes": 49,
					"arxiv_id": "2608.23041",
					"first_author": "Sungho Park",
					"org": "Microsoft"
				},
				{
					"rank": 3,
					"title": "On-Policy Self-Distillation in Diffusion Models",
					"title_zh": "扩散模型中的同策略自蒸馏",
					"title_ja": "拡散モデルにおけるオンポリシー自己蒸留",
					"url": "https://huggingface.co/papers/2608.24646",
					"summary": "ByteDance Seed introduces DiffusionOPSD, which converts image-level reward guidance into explicit targets for a diffusion model's clean-output predictions. A frozen behavior policy supplies the trajectories and anchors, and reward gradients build bounded positive and negative targets around each anchor. The point is to specify how an intermediate denoising prediction should change, which endpoint rewards do not.",
					"summary_zh": "字节跳动 Seed 团队提出 DiffusionOPSD，将图像级的奖励引导转化为扩散模型在采样查询状态下对干净输出预测的显式目标。冻结的行为策略提供轨迹与锚点，奖励梯度则在每个锚点周围构造有界的正负目标。其要点在于给出中间去噪预测该如何调整，而终点奖励无法提供这一信息。",
					"summary_ja": "ByteDance Seed は DiffusionOPSD を提案した。画像レベルの報酬による誘導を、サンプリングしたクエリ状態における拡散モデルのクリーン出力予測に対する明示的な目標へと変換する。凍結した行動方策が軌跡とアンカーを供給し、報酬勾配が各アンカーの周囲に有界の正負の目標を作る。狙いは、終点の報酬では示せない「途中のノイズ除去予測をどう変えるべきか」を与えることにある。",
					"upvotes": 47,
					"arxiv_id": "2608.24646",
					"first_author": "Wei Zhou",
					"org": "ByteDance Seed"
				},
				{
					"rank": 4,
					"title": "SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation",
					"title_zh": "SecOPD：以同策略蒸馏缓解自适应提示注入",
					"title_ja": "SecOPD: オンポリシー蒸留による適応的プロンプトインジェクションの緩和",
					"url": "https://huggingface.co/papers/2608.21500",
					"summary": "The authors argue that defensive finetuning still fails against adaptive prompt injection because DPO and GRPO grade an entire output with one sequence-level signal, treating every token alike. SecOPD uses on-policy distillation to supply token-level feedback instead. Prompt injection, in which text hidden in fetched pages, files or email redirects an agent, is listed as the top threat to AI agents.",
					"summary_zh": "作者认为，防御性微调之所以仍难以抵御自适应提示注入，是因为 DPO 与 GRPO 用单一的序列级信号评判整段输出，对每个 token 一视同仁。SecOPD 改用同策略蒸馏来提供 token 级反馈。提示注入指攻击者把指令藏进智能体抓取的网页、文件或邮件中以劫持其行为，被列为 AI 智能体面临的头号威胁。",
					"summary_ja": "著者らは、防御的なファインチューニングが適応的プロンプトインジェクションに依然として破られるのは、DPO や GRPO が出力全体を一つのシーケンス単位の信号で評価し、すべてのトークンを同等に扱うためだと論じる。SecOPD は代わりにオンポリシー蒸留でトークン単位のフィードバックを与える。取得したページやファイル、メールに仕込まれた指示でエージェントを乗っ取るプロンプトインジェクションは、AI エージェントに対する最大の脅威に挙げられている。",
					"upvotes": 36,
					"arxiv_id": "2608.21500",
					"first_author": "Yibo Peng"
				},
				{
					"rank": 5,
					"title": "The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models",
					"title_zh": "掩码不等于模型：审计注意力、状态空间与混合序列模型中的前缀不变性",
					"title_ja": "マスクはモデルではない: アテンション・状態空間・ハイブリッド系列モデルにおける前置不変性の監査",
					"url": "https://huggingface.co/papers/2608.22876",
					"summary": "The paper formalizes prefix invariance - a representation at position t must not depend on later inputs - and gives a two-forward-pass audit, with no training or gradients, that localizes where causality breaks. Mask inspection is not enough: leaks can travel through scans or normalization. Across 192 injected-fault trials it caught none, while the audit found all 192 plus a defect in Zamba2 and Nemotron-H.",
					"summary_zh": "论文形式化了前缀不变性，即位置 t 上的表示不得依赖其后的输入，并给出一种只需两次前向传播、无需训练或梯度的审计方法，可定位因果性在何处被打破。仅检查注意力掩码并不够，泄漏可能经由扫描或归一化发生。在 192 组注入故障的试验中，掩码检查一个都没发现，而该审计方法全部定位，并另外查出 Zamba2 与 Nemotron-H 中的缺陷。",
					"summary_ja": "本論文は前置不変性（位置 t の表現が後続の入力に依存してはならないこと）を定式化し、学習も勾配も不要で順伝播 2 回だけで因果性の破れ箇所を特定する監査手法を示した。アテンション・マスクの点検だけでは不十分で、スキャンや正規化を経由して漏れが生じうる。故障を注入した 192 試行で、マスク点検は 1 件も検出できなかったのに対し、この監査は全件を特定し、さらに Zamba2 と Nemotron-H の欠陥も見つけた。",
					"upvotes": 29,
					"arxiv_id": "2608.22876",
					"first_author": "Taebong Kim",
					"org": "VIDraft"
				}
			],
			"updated_at": "2026-08-26 13:05 PDT"
		},
		{
			"date": "2026-08-25",
			"items": [
				{
					"rank": 1,
					"title": "TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming",
					"title_zh": "TLive-Omni：面向电商直播的全模态理解模型",
					"title_ja": "TLive-Omni: EC ライブ配信向けのオムニモーダル理解モデル",
					"url": "https://huggingface.co/papers/2608.20958",
					"arxiv_id": "2608.20958",
					"upvotes": 52,
					"first_author": "Yibo Hu",
					"org": "TaoLive AIGC",
					"summary": "TLive-Omni is an omni-modal model built for e-commerce live streams, where product facts are scattered across speech, video frames, product images, overlaid text and viewer questions. It maps image, video, audio and text into a single representation space, and adds Per-vGrid, a timestamped token layout that pairs each video grid with its matching audio inside explicit boundary tokens.",
					"summary_zh": "TLive-Omni 是一款面向电商直播的全模态理解模型，直播中的商品信息分散在语音、视频画面、商品图、叠加文字和观众提问之中。该模型将图像、视频、音频与文本映射到统一表示空间，并提出带时间戳的 token 组织方式 Per-vGrid，在显式边界标记内把每个视频网格与对应音频配对。",
					"summary_ja": "TLive-Omni は EC ライブ配信向けのオムニモーダル理解モデルで、商品情報が音声、映像フレーム、商品画像、重畳テキスト、視聴者の質問に分散する状況を対象とする。画像・映像・音声・テキストを単一の表現空間に写像し、各映像グリッドと対応する音声を明示的な境界トークン内で対応付けるタイムスタンプ付きトークン構成 Per-vGrid を導入した。"
				},
				{
					"rank": 2,
					"title": "InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter",
					"title_zh": "InfinityEdit：用轻量 Edit-Ignition 适配器实现无限视频编辑",
					"title_ja": "InfinityEdit: 軽量な Edit-Ignition アダプタによる無限動画編集",
					"url": "https://huggingface.co/papers/2608.20910",
					"arxiv_id": "2608.20910",
					"upvotes": 35,
					"first_author": "Yunze Tong",
					"summary": "Instruction-based video editing usually assumes an in-place edit, aligning the output frame by frame with a fixed source clip. InfinityEdit takes on what the authors call infinite video editing, where an instruction given on a preceding segment must carry into frames that arrive later, as when restyling a live game or extending a camera move on an ongoing shot.",
					"summary_zh": "基于指令的视频编辑通常假设原位编辑，即输出与固定的源片段逐帧对齐。InfinityEdit 研究的是作者所称的无限视频编辑：针对前一段画面给出的指令，必须延续到随后陆续到来的帧上，例如为直播游戏改变画风，或为正在进行的镜头延展运镜。",
					"summary_ja": "指示ベースの動画編集は通常、出力を固定のソース映像とフレーム単位で対応させるインプレース編集を前提とする。InfinityEdit は著者らが無限動画編集と呼ぶ設定を扱い、先行区間に与えた指示を、後から到着するフレームにも引き継ぐ。ライブゲームの画風変更や、進行中のショットへのカメラワーク付与などが対象となる。"
				},
				{
					"rank": 3,
					"title": "MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks",
					"title_zh": "MobilePA-Bench：在复杂真实任务上评测移动端规划智能体",
					"title_ja": "MobilePA-Bench: 複雑な実世界タスクにおけるモバイル・プランナーエージェントの評価",
					"url": "https://huggingface.co/papers/2608.23035",
					"arxiv_id": "2608.23035",
					"upvotes": 33,
					"first_author": "Yi Zhu",
					"org": "Tongyi-MAI",
					"summary": "MobilePA-Bench is an interactive, stateful, tool-centric benchmark for on-device LLM agents acting as phone copilots. The authors argue existing tests fall into two camps with matching blind spots: GUI benchmarks that measure only surface screen manipulation, and static function-calling sets that match APIs offline, away from real runtime constraints.",
					"summary_zh": "MobilePA-Bench 是一个面向端侧 LLM 智能体的交互式、有状态、以工具为中心的评测基准，考察其作为手机助手的能力。作者认为现有评测分为两类且各有盲区：GUI 基准只衡量表层的屏幕操作，而静态函数调用测试则以离线 API 匹配为准，脱离真实运行时约束。",
					"summary_ja": "MobilePA-Bench は、スマートフォン上のコパイロットとして動作するオンデバイス LLM エージェント向けの、対話的で状態を持つツール中心のベンチマークである。著者らは既存の評価が二分され、GUI ベンチマークは表層的な画面操作しか測らず、静的な関数呼び出しの評価はオフラインの API 一致に依存して実行時の制約から乖離していると指摘する。"
				},
				{
					"rank": 4,
					"title": "Prime Agent: A Self-Improving RLM Harness",
					"title_zh": "Prime Agent：可自我改进的 RLM 智能体框架",
					"title_ja": "Prime Agent: 自己改善する RLM ハーネス",
					"url": "https://huggingface.co/papers/2608.23552",
					"arxiv_id": "2608.23552",
					"upvotes": 32,
					"first_author": "Seth Karten",
					"org": "Prime Intellect",
					"summary": "Prime Intellect describes Prime Agent, an open-source harness for long-horizon evaluation and coding-agent workflows. A persistent IPython REPL implements the Recursive Language Model abstraction for programmatic context processing and test-time compute, while a Continual Harness carries histories, memories, skills, prompts and subagent specifications across trajectories.",
					"summary_zh": "Prime Intellect 介绍了开源框架 Prime Agent，面向长程任务评测与编码智能体工作流。其常驻的 IPython REPL 实现了递归语言模型（RLM）抽象，用于程序化的上下文处理与测试时计算；Continual Harness 则在多条轨迹之间保留历史、记忆、技能、提示词与子智能体配置。",
					"summary_ja": "Prime Intellect は、長期タスクの評価とコーディングエージェントのワークフローに向けたオープンソースのハーネス Prime Agent を発表した。常駐する IPython REPL が再帰言語モデル（RLM）の抽象を実装し、プログラム的な文脈処理とテスト時計算を担う。Continual Harness は履歴、記憶、スキル、プロンプト、サブエージェント定義を軌跡をまたいで保持する。"
				},
				{
					"rank": 5,
					"title": "Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion",
					"title_zh": "Block3D：基于分块扩散的高效文本到 3D 生成",
					"title_ja": "Block3D: ブロック単位の拡散による効率的なテキストから 3D 生成",
					"url": "https://huggingface.co/papers/2608.19567",
					"arxiv_id": "2608.19567",
					"upvotes": 25,
					"first_author": "Bowen Cui",
					"org": "Zhejiang University",
					"summary": "Block3D generates 3D shapes from text using a block-wise diffusion framework that partitions the discrete shape representation instead of decoding tokens one at a time or refining the whole representation at every step. The authors note that autoregressive decoding cannot revise its own errors, while diffusion and flow models pay full-representation cost each iteration.",
					"summary_zh": "Block3D 采用分块扩散框架从文本生成 3D 形状，对离散形状表示进行分块处理，而非逐个自回归解码 token 或在每一步刷新整体表示。作者指出，自回归解码无法修正自身错误，而扩散与流匹配模型每次迭代都要付出处理完整表示的代价。",
					"summary_ja": "Block3D は、離散的な形状表現をブロックに分割する拡散フレームワークでテキストから 3D 形状を生成する。トークンを逐次デコードしたり、毎ステップで表現全体を更新したりする方式とは異なる。著者らは、自己回帰デコードは自らの誤りを修正できず、拡散やフローの手法は反復ごとに表現全体を処理するコストを払うと指摘する。"
				}
			],
			"updated_at": "2026-08-25 13:05 PDT"
		},
		{
			"date": "2026-08-24",
			"items": [
				{
					"rank": 1,
					"title": "ParaTempo: Efficient Parallel Reasoning via Temporal Confidence",
					"title_zh": "ParaTempo：基于时间置信度的高效并行推理",
					"title_ja": "ParaTempo: 時間的確信度による効率的な並列推論",
					"url": "https://huggingface.co/papers/2608.16425",
					"arxiv_id": "2608.16425",
					"upvotes": 26,
					"first_author": "Xuteng Zhang",
					"org": "Shanghai Jiao Tong University",
					"summary": "ParaTempo is a training-free framework for asynchronous parallel reasoning in large reasoning models. The authors argue that existing controls for parallel branches - final-answer consensus, local token confidence, or isolated intermediate probes - are delayed, weakly tied to reasoning progress, or too noisy for branch-level decisions, and propose a temporal confidence signal instead.",
					"summary_zh": "ParaTempo 是面向大型推理模型的免训练异步并行推理框架。作者指出，现有的并行分支控制手段（最终答案共识、局部 token 置信度或孤立的中间探针）存在信号滞后、与推理进展关联薄弱、或噪声过大难以支撑分支级决策等问题，转而提出以时间置信度作为控制信号。",
					"summary_ja": "ParaTempo は、大規模推論モデル向けの学習不要な非同期並列推論フレームワークである。著者らは、最終回答の多数決、局所的なトークン確信度、孤立した中間プローブといった既存の分岐制御手法は、信号が遅れる、推論の進捗との結びつきが弱い、あるいはノイズが多く分岐単位の制御には使えないと指摘し、代わりに時間的確信度を用いる手法を示した。"
				},
				{
					"rank": 2,
					"title": "Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs",
					"title_zh": "超越正确性：混合思考多模态大模型的响应行为评测与对齐",
					"title_ja": "正解率を超えて: ハイブリッド思考型 MLLM における応答挙動のベンチマークとアライメント",
					"url": "https://huggingface.co/papers/2608.12781",
					"arxiv_id": "2608.12781",
					"upvotes": 16,
					"first_author": "Xinming Wang",
					"org": "Tencent",
					"summary": "Tencent researchers examine multimodal models that switch between deliberative thinking and low-latency non-thinking inference, arguing that both modes should meet the same user-facing standard. The work evaluates response-pattern failures alongside task accuracy and tests whether the two interfaces preserve acceptable final-response behavior.",
					"summary_zh": "腾讯的研究者考察了可在深思模式与低延迟非思考模式之间切换的多模态模型，认为两种模式面向用户时应达到同一标准。研究在任务准确率之外，同时评估响应模式层面的失败，并检验两种接口是否都能保持可接受的最终响应行为。",
					"summary_ja": "テンセントの研究チームは、熟考モードと低遅延の非思考モードを切り替えるマルチモーダルモデルを対象に、どちらのモードもユーザーに対して同じ水準を満たすべきだと論じる。タスク正解率に加えて応答パターンの失敗を評価し、二つのインターフェースが許容できる最終応答の挙動を保てているかを検証した。"
				},
				{
					"rank": 3,
					"title": "EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking",
					"title_zh": "EviRank：面向多模态图像重排序的结构化相关性证据",
					"title_ja": "EviRank: マルチモーダル画像リランキングのための構造化された関連性エビデンス",
					"url": "https://huggingface.co/papers/2608.20886",
					"arxiv_id": "2608.20886",
					"upvotes": 10,
					"first_author": "Enjun Du",
					"summary": "EviRank recasts multimodal image re-ranking as a semantic constraint satisfaction problem, parsing text-only, image-only or composed queries into structured relevance evidence. The authors argue existing re-rankers either compress multifaceted relevance into an opaque embedding or lean on free-form chain-of-thought that drops or hallucinates fine-grained constraints.",
					"summary_zh": "EviRank 将多模态图像重排序重新表述为语义约束满足问题，把纯文本、纯图像或组合式查询解析为结构化的相关性证据。作者认为，现有重排序方法要么把多维度的相关性压缩进不透明的向量表示，要么依赖自由形式的思维链，容易遗漏或臆造细粒度约束。",
					"summary_ja": "EviRank は、マルチモーダル画像のリランキングを意味的な制約充足問題として捉え直し、テキストのみ、画像のみ、あるいは複合的なクエリを構造化された関連性エビデンスへと解析する。著者らは、既存のリランカーが多面的な関連性を不透明な埋め込みに圧縮するか、細かな制約を取りこぼしたり捏造したりしやすい自由記述の思考連鎖に依存していると指摘する。"
				},
				{
					"rank": 4,
					"title": "UniSpace: Unified Visual Representation and Scalable Multimodal Modeling",
					"title_zh": "UniSpace：统一视觉表征与可扩展多模态建模",
					"title_ja": "UniSpace: 統一的な視覚表現とスケーラブルなマルチモーダルモデリング",
					"url": "https://huggingface.co/papers/2608.08676",
					"arxiv_id": "2608.08676",
					"upvotes": 8,
					"first_author": "Jinbo Yan",
					"org": "LongCat",
					"summary": "UniSpace asks whether understanding, generation and editing can share one visual representation space built from a pretrained semantic vision transformer. The authors note that final tokens from semantic encoders discard fine detail and hurt pixel reconstruction, and show that frozen ViT blocks are not inherently unable to preserve it.",
					"summary_zh": "UniSpace 探讨能否让理解、生成与编辑共享一个由预训练语义视觉 Transformer 构建的视觉表征空间。作者指出，语义编码器的最终层 token 丢弃了细粒度细节、损害像素级重建，并表明冻结的 ViT 模块本身并非无法保留这些细节。",
					"summary_ja": "UniSpace は、理解・生成・編集を、事前学習済みの意味的 Vision Transformer から構築した単一の視覚表現空間で扱えるかを問う。著者らは、意味エンコーダーの最終トークンが細部を捨ててしまい画素単位の再構成を損なうと指摘したうえで、凍結した ViT ブロックが本質的に細部を保持できないわけではないことを示した。"
				},
				{
					"rank": 5,
					"title": "Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models",
					"title_zh": "划分支撑集，重建残差：面向视频生成与世界模型的免训练稀疏注意力",
					"title_ja": "サポートを分割し残差を再構成する: 動画生成と世界モデルのための学習不要なスパースアテンション",
					"url": "https://huggingface.co/papers/2608.18484",
					"arxiv_id": "2608.18484",
					"upvotes": 6,
					"first_author": "Pardis Taghavi",
					"org": "Texas A&M University",
					"summary": "The paper introduces SparsePR, a training-free block-sparse attention method for video transformers that pairs response-coupled partitioning with probe-fitted residual reconstruction. The authors argue that row-wise attention concentration alone does not specify a usable sparse operator, because queries sharing a block route may have poorly overlapping supports.",
					"summary_zh": "论文提出 SparsePR，一种面向视频 Transformer 的免训练块稀疏注意力方法，将响应耦合的分区与探针拟合的残差重建结合起来。作者指出，仅有按行的注意力集中度并不足以确定可用的稀疏算子，因为共享同一分块路由的查询，其支撑集可能重叠很少。",
					"summary_ja": "本論文は、動画向け Transformer のための学習不要なブロックスパースアテンション手法 SparsePR を提案する。応答結合型の分割と、プローブで当てはめた残差再構成を組み合わせる点が特徴だ。著者らは、行単位のアテンション集中度だけでは実行可能なスパース演算子は定まらないと述べる。同じブロック経路を共有するクエリでも、サポートがほとんど重ならない場合があるためだ。"
				}
			],
			"updated_at": "2026-08-24 13:04 PDT"
		},
		{
			"date": "2026-08-22",
			"items": [
				{
					"rank": 1,
					"title": "Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See",
					"title_zh": "用低资源语言思考：SFT 建立了什么，RL 修正了什么，准确率看不见什么",
					"title_ja": "低資源言語で考える: SFT が築くもの、RL が直すもの、精度が見られないもの",
					"url": "https://huggingface.co/papers/2608.17744",
					"arxiv_id": "2608.17744",
					"upvotes": 13,
					"first_author": "Ayoub Kirouane",
					"org": "KIEFER",
					"summary": "Three frontier mixture-of-experts models with 3.6 to 4.0B active parameters were fine-tuned to reason in Greek, and accuracy barely moved. Changing only the random seed shifted scores by 7.7 points, more than any data or recipe effect measured, making the benchmark itself noise at that scale. What did change was the language of thought: base models produced Greek reasoning in 0 of 1,000 traces.",
					"summary_zh": "研究者对三个激活参数为 36 亿至 40 亿的前沿混合专家模型进行微调，使其用希腊语推理，结果准确率几乎没有变化。仅更换随机种子就能使分数波动 7.7 分，超过所测的任何数据或配方带来的影响，说明该规模下基准本身就是噪声。真正改变的是思考所用的语言：基座模型在 1,000 条推理轨迹中用希腊语思考的次数为 0。",
					"summary_ja": "アクティブパラメータ 36 億〜40 億の最先端 MoE モデル 3 つをギリシャ語で推論するようファインチューニングしたが、精度はほとんど動かなかった。乱数シードを変えるだけでスコアは 7.7 ポイント動き、測定したどのデータ・レシピの効果より大きい。変化したのは思考の言語で、ベースモデルは 1,000 件の推論軌跡のうち 0 件しかギリシャ語で考えていなかった。"
				},
				{
					"rank": 2,
					"title": "Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses",
					"title_zh": "层级式自我改进：面向任务的可演化代理运行框架",
					"title_ja": "階層的自己改善: タスク特化で進化するエージェントハーネスの枠組み",
					"url": "https://huggingface.co/papers/2608.08466",
					"arxiv_id": "2608.08466",
					"upvotes": 10,
					"first_author": "Tailin Zhou",
					"org": "HKUST",
					"summary": "LLM agents are usually improved by editing prompts, tools or workflows, while the harness that executes around the model stays fixed after deployment. This work makes the harness task-specific and continuously evolvable: each task family keeps its own harness, hot-swapped across iterations through a fixed task-injection seam and rewritten from environment feedback, with a single frozen model throughout.",
					"summary_zh": "大模型代理通常靠修改提示词、工具或工作流来改进，而围绕模型执行的运行框架（harness）在部署后往往固定不变。该工作让运行框架按任务定制并持续演化：每类任务维护自己的框架，通过固定的任务注入接口在迭代中热替换，并依据环境反馈重写，整个过程只用一个冻结的模型。",
					"summary_ja": "LLM エージェントの改善は通常プロンプトやツール、ワークフローの手直しで行われ、モデルの周囲で実行を担うハーネスは配備後に固定されたままになる。本研究はハーネスをタスク特化かつ継続的に進化するものとし、タスク系統ごとに専用ハーネスを保持して、固定の注入インターフェース経由で反復ごとにホットスワップし、環境フィードバックから書き換える。モデルは 1 つを凍結したまま用いる。"
				},
				{
					"rank": 3,
					"title": "EXIMO: VLM Guided Exploration of VLA Policies",
					"title_zh": "EXIMO：由视觉语言模型引导的 VLA 策略探索",
					"title_ja": "EXIMO: VLM が導く VLA ポリシーの探索",
					"url": "https://huggingface.co/papers/2608.19891",
					"arxiv_id": "2608.19891",
					"upvotes": 10,
					"first_author": "Bhavya Sukhija",
					"org": "Deepmind",
					"summary": "DeepMind takes on how to finetune robot policies for new tasks on the fly. State-of-the-art manipulation rests on behavior cloning of billion-parameter vision-language-action models trained on huge teleoperation datasets, and adapting them is still open: more teleoperation costs hundreds of hours of human labor, while reinforcement learning is sample-hungry. EXIMO has a vision-language model guide exploration.",
					"summary_zh": "DeepMind 研究如何让机器人策略即时适配新任务。当前最先进的操作能力依赖对十亿参数级视觉-语言-动作模型做行为克隆，训练数据来自海量遥操作记录，而微调仍是未解问题：继续采集遥操作数据需要数百小时人力，强化学习又极其耗样本。EXIMO 改由视觉语言模型引导探索。",
					"summary_ja": "DeepMind は、ロボットのポリシーを新しいタスクへその場で適応させる問題に取り組む。最先端の操作性能は、膨大な遠隔操作データで学習した数十億パラメータの視覚-言語-行動モデルの模倣学習に依存し、その微調整は未解決だ。遠隔操作データの追加収集は数百時間の人手を要し、強化学習はサンプル効率が悪い。EXIMO は代わりに視覚言語モデルに探索を導かせる。"
				},
				{
					"rank": 4,
					"title": "Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization",
					"title_zh": "注入、对齐、恢复：面向免检索文档知识内化的分阶段后训练",
					"title_ja": "Inject, Align, Recover: 検索不要の文書知識内在化に向けた段階的ポストトレーニング",
					"url": "https://huggingface.co/papers/2608.20281",
					"arxiv_id": "2608.20281",
					"upvotes": 10,
					"first_author": "Qian Kou",
					"org": "Beijing Academy of Artificial Intelligence",
					"summary": "Language models often fail on questions about a bounded document collection when the sources are not retrieved at inference time. IAR is a three-stage post-training framework that separates the problem: Inject turns source documents into continuation data, Align tunes question-answering behavior, and Recover restores general ability. The authors present it as an alternative to conventional continued pretraining.",
					"summary_zh": "当推理时不检索原文，语言模型常常答不出关于某个有限文档集合的问题。IAR 是一个三阶段后训练框架，把问题拆开处理：Inject 将源文档转换为续写数据，Align 调整问答行为，Recover 恢复通用能力。作者将其定位为常规继续预训练的替代方案。",
					"summary_ja": "推論時に原文を検索しない場合、言語モデルは限定された文書集合に関する質問にしばしば答えられない。IAR は問題を三つの段階に分ける後訓練の枠組みで、Inject が原文を継続予測用データに変換し、Align が質問応答の振る舞いを整え、Recover が汎用能力を回復させる。著者らは従来の継続事前学習に代わる手法として提示している。"
				},
				{
					"rank": 5,
					"title": "NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video",
					"title_zh": "NARU：日语超长视频中叙事演变与文化细微差别理解的基准",
					"title_ja": "NARU: 日本語の超長尺動画における物語の展開と文化的機微の理解を測るベンチマーク",
					"url": "https://huggingface.co/papers/2608.13210",
					"arxiv_id": "2608.13210",
					"upvotes": 8,
					"first_author": "Yuheng Huang",
					"org": "The University of Tokyo",
					"summary": "NARU is a benchmark for tracking an evolving narrative and reading implicit social meaning in Japanese long-form video. It holds 1,481 questions grounded in 155 videos totaling 146.8 hours, spanning four narrative and five cultural dimensions. Existing benchmarks rarely test those two capabilities together, especially in high-context non-English media.",
					"summary_zh": "NARU 是一个基准，用于考察在日语长视频中追踪叙事发展、并读出未言明的社会含义的能力。它包含 1,481 道题，取材自 155 段、合计 146.8 小时的视频，覆盖四个叙事维度和五个文化维度。现有基准很少同时评估这两类能力，在高语境的非英语媒体中尤其如此。",
					"summary_ja": "NARU は、日本語の長尺動画において物語の展開を追い、明示されない社会的な意味を読み取る能力を測るベンチマーク。155 本・計 146.8 時間の動画に基づく 1,481 問からなり、物語の 4 次元と文化の 5 次元をカバーする。既存のベンチマークでこの二つを同時に評価するものは少なく、高文脈かつ非英語のメディアでは特に乏しい。"
				}
			],
			"updated_at": "2026-08-22 13:04 PDT"
		},
		{
			"date": "2026-08-21",
			"items": [
				{
					"rank": 1,
					"title": "FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis",
					"title_zh": "FACET：在终端任务合成中保留源意图与可执行状态",
					"title_ja": "FACET: ターミナルタスク合成における元の意図と実行可能な状態の保持",
					"url": "https://huggingface.co/papers/2608.18580",
					"arxiv_id": "2608.18580",
					"upvotes": 102,
					"first_author": "Kou Shi",
					"org": "University of Science and Technology of China",
					"summary": "Training terminal agents needs executable tasks at scale, yet each task couples an instruction, an initialized environment, a reference solution and a verifier, and inconsistent assumptions across those parts leave tasks unsolvable or wrongly graded. FACET synthesizes tasks while carrying the goals, dependencies and state transitions of the original sources through the pipeline.",
					"summary_zh": "训练终端智能体需要大规模可执行任务，但每个任务同时包含指令、初始化环境、参考解法与验证器，一旦各部分的假设互不一致，任务就会无解或被错误评判。FACET 在合成过程中保留原始素材中的目标、依赖关系与状态转移，以细粒度的智能体化方式构建任务。",
					"summary_ja": "ターミナル操作を担うエージェントの学習には実行可能なタスクが大量に要るが、各タスクは指示、初期化された環境、参照解、検証器を同時に抱えており、前提が食い違えばタスクは解けないか誤って採点される。FACET は元の素材にある目標、依存関係、状態遷移を合成の各段階に持ち越してタスクを構築する。"
				},
				{
					"rank": 2,
					"title": "SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?",
					"title_zh": "SWE-bench Science：编码智能体能否解决科学领域的工程任务？",
					"title_ja": "SWE-bench Science: コーディングエージェントは科学の工学的課題を解けるか",
					"url": "https://huggingface.co/papers/2608.19799",
					"arxiv_id": "2608.19799",
					"upvotes": 51,
					"first_author": "Zhipeng Xu",
					"org": "OpenMOSS",
					"summary": "OpenMOSS introduced SWE-bench Science, a repository-level benchmark of 119 tasks drawn from 98 GitHub repositories across 20 scientific domains. It is built to show why coding agents fail when repairing scientific software rather than only whether they succeed, on the argument that faulty scientific code can compromise the evidence behind published conclusions.",
					"summary_zh": "OpenMOSS 推出 SWE-bench Science，这是一个仓库级基准，包含取自 20 个科学领域、98 个 GitHub 仓库的 119 项任务。它的设计不只看编码智能体能否修复科学软件，更要揭示它们失败的原因，理由是有缺陷的科学代码可能动摇论文结论所依赖的证据。",
					"summary_ja": "OpenMOSS は SWE-bench Science を公開した。20 の科学分野、98 の GitHub リポジトリから集めた 119 課題からなるリポジトリ規模のベンチマークである。成否の集計にとどまらず、コーディングエージェントが科学ソフトの修復でなぜ失敗するのかを示す設計で、欠陥のあるコードは論文の結論を支える証拠自体を損ないうるという問題意識に立つ。"
				},
				{
					"rank": 3,
					"title": "WithEveryone: Unified Planning and Identity Grounding for Group Image Generation",
					"title_zh": "WithEveryone：面向合影生成的统一规划与身份对齐",
					"title_ja": "WithEveryone: 集合写真生成のための統一的な計画立案と同一性の接地",
					"url": "https://huggingface.co/papers/2608.20336",
					"arxiv_id": "2608.20336",
					"upvotes": 34,
					"first_author": "Hengyuan Xu",
					"org": "Tencent Hunyuan",
					"summary": "Tencent Hunyuan's WithEveryone generates group images holding up to ten specified people without their identities blurring together. Each reference is injected as an addressed token, the model predicts a structured identity and layout plan, and that plan is rendered back as a visual condition for generation.",
					"summary_zh": "腾讯混元提出 WithEveryone，可生成最多包含十个指定人物的合影，而不至于让各人身份彼此混淆。每个参考人物以带地址的 token 注入，模型先预测结构化的身份与布局方案，再把该方案渲染为视觉条件用于生成。",
					"summary_ja": "テンセント混元の WithEveryone は、指定した最大 10 人を含む集合写真を、各人の同一性が混ざらないように生成する。参照人物ごとに宛先付きのトークンを注入し、構造化された同一性とレイアウトの計画を予測したうえで、その計画を視覚的な条件として描き戻す仕組みである。"
				},
				{
					"rank": 4,
					"title": "MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use",
					"title_zh": "MemTrapBench：大模型记忆使用中的认知陷阱基准",
					"title_ja": "MemTrapBench: LLM の記憶利用に潜む認知の罠を測るベンチマーク",
					"url": "https://huggingface.co/papers/2608.20202",
					"arxiv_id": "2608.20202",
					"upvotes": 24,
					"first_author": "Mengru Wang",
					"org": "ZJUNLP",
					"summary": "MemTrapBench targets what its authors call memory-induced cognitive traps: even a faithfully recorded and relevant memory can distort a model's reasoning and hurt performance on the task at hand. Existing memory benchmarks mostly check whether information was extracted, stored and retrieved correctly, not how the retrieved text reshapes the answer.",
					"summary_zh": "MemTrapBench 针对作者所称的记忆诱发认知陷阱：即便是被如实记录且语义相关的记忆，也可能扭曲模型推理，拖累当前任务的表现。现有记忆基准大多只检验信息是否被正确抽取、存储与检索，而忽略了取回的内容如何改变最终作答。",
					"summary_ja": "MemTrapBench は著者が記憶由来の認知の罠と呼ぶ現象を対象とする。忠実に記録され意味的にも関連する記憶であっても、モデルの推論をゆがめ、目の前の課題の成績を落としうるという問題である。既存の記憶ベンチマークは情報が正しく抽出・保存・検索されたかを見るにとどまり、取り出した記憶が答えをどう作り替えるかは扱ってこなかった。"
				},
				{
					"rank": 5,
					"title": "SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback",
					"title_zh": "SkillEvo：从多轮交互反馈中自我更新的进化梯度",
					"title_ja": "SkillEvo: マルチターンの対話フィードバックから自己更新する進化勾配",
					"url": "https://huggingface.co/papers/2608.13120",
					"arxiv_id": "2608.13120",
					"upvotes": 21,
					"first_author": "Qianxi Yan",
					"org": "Tencent",
					"summary": "Agent skills are usually hand-authored or produced in a single generation pass, so they never learn from the failures they cause. Tencent's SkillEvo argues that recent feedback loops stall because they score skills on single-turn question answering, which hides defects that only surface across several turns, and instead draws its evolution signal from multi-turn interaction.",
					"summary_zh": "智能体技能通常靠人工撰写，或由模型一次性生成，因而无法从自己引发的失败中学习。腾讯提出的 SkillEvo 认为，已有的反馈闭环之所以很快停滞，是因为它们以单轮问答来评估技能，掩盖了只有在多轮交互中才暴露的缺陷，因此改从多轮交互反馈中提取进化信号。",
					"summary_ja": "エージェントのスキルは人手で書かれるか、モデルが一度の生成で作るのが常で、自らが招いた失敗から学ぶ経路を持たない。テンセントの SkillEvo は、既存のフィードバック閉ループが早々に頭打ちになるのは単一ターンの質疑応答でスキルを評価するためであり、それでは複数ターンでしか現れない欠陥が見えないとして、多ターンの対話から進化の信号を取り出す。"
				}
			],
			"updated_at": "2026-08-21 13:08 PDT"
		},
		{
			"date": "2026-08-20",
			"items": [
				{
					"rank": 1,
					"title": "SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation",
					"title_zh": "SemComp-Bench：视频生成中语义任务完成度的基准测试",
					"title_ja": "SemComp-Bench: 動画生成における意味的タスク達成度のベンチマーク",
					"url": "https://huggingface.co/papers/2608.17426",
					"upvotes": 151,
					"arxiv_id": "2608.17426",
					"first_author": "Keyu Tu",
					"org": "FrameX-AI",
					"summary": "SemComp-Bench recasts video generation as an outcome-oriented task: a clip counts as a success only if it reaches the intended result and stays semantically grounded in the reference image. Evaluation judges the generated outcome rather than demanding a complete sequence of intermediate steps or conventional appearance consistency. It is the most upvoted paper on Hugging Face today with 151 votes.",
					"summary_zh": "SemComp-Bench 把视频生成重新定义为面向结果的任务：只有当生成片段达成预期结果、并在语义上与参考图像保持对应时才算成功。评测只考察生成的最终结果，而不要求呈现完整的中间步骤序列或传统意义上的外观一致性。该论文以 151 票成为今天 Hugging Face 上得票最高的一篇。",
					"summary_ja": "SemComp-Bench は動画生成を結果志向のタスクとして捉え直す。意図した成果に到達し、かつ参照画像と意味的に対応している場合にのみ成功とみなす。評価は生成された成果そのものを見るもので、中間手順の完全な系列や従来型の見た目の一貫性は求めない。本日の Hugging Face で最多となる 151 票を集めた。"
				},
				{
					"rank": 2,
					"title": "OmniScientist: An Omni-Modal Omni-Discipline AI Scientist",
					"title_zh": "OmniScientist：全模态、跨学科的 AI 科学家",
					"title_ja": "OmniScientist: 全モーダル・全分野に対応する AI サイエンティスト",
					"url": "https://huggingface.co/papers/2608.13558",
					"upvotes": 83,
					"arxiv_id": "2608.13558",
					"first_author": "Bobo Li",
					"org": "National University of Singapore",
					"summary": "OmniScientist is an end-to-end AI scientist that runs multidisciplinary research across modalities instead of reasoning only over text, code, labels or precomputed summaries. The authors argue that existing systems leave out the spatial, temporal, cross-channel and procedural relations that scientific discovery actually turns on, so workflow coverage alone is not enough.",
					"summary_zh": "OmniScientist 是一个端到端的 AI 科学家系统，跨模态开展多学科研究，而不是仅在文本、代码、标签或预先汇总的结果上做推理。作者认为现有系统遗漏了科学发现真正依赖的空间、时间、跨通道与流程性关系，因此仅覆盖完整工作流并不足够。",
					"summary_ja": "OmniScientist は、テキストやコード、ラベル、事前に要約された情報だけで推論するのではなく、複数のモダリティにまたがって多分野の研究を行うエンドツーエンドの AI サイエンティストである。著者らは、既存システムが科学的発見の鍵となる空間的、時間的、チャネル横断的、手続き的な関係を取りこぼしており、ワークフローを網羅するだけでは不十分だと主張する。"
				},
				{
					"rank": 3,
					"title": "SPADE: Self-Play in Adaptive Synthetic Executable Environments",
					"title_zh": "SPADE：在自适应合成可执行环境中的自博弈",
					"title_ja": "SPADE: 適応的な合成実行環境における自己対戦",
					"url": "https://huggingface.co/papers/2608.19197",
					"upvotes": 40,
					"arxiv_id": "2608.19197",
					"first_author": "Bo Liu",
					"org": "spade-rl",
					"summary": "SPADE is a self-play reinforcement learning framework in which one language model takes two roles: an Environment Designer that writes long-horizon training environments as executable code behind a Gym-style reset and step interface, and a Reasoning Agent that learns inside them. The aim is a goal distribution that keeps expanding as the learner scales, which hand-curated or frozen environment pools cannot do.",
					"summary_zh": "SPADE 是一个自博弈强化学习框架，由同一个语言模型扮演两种角色：环境设计者把长程训练环境写成可执行代码，并暴露 Gym 风格的 reset 与 step 接口；推理智能体则在其中学习。其目标是让任务分布随学习者一同扩展，而人工整理或固定不变的环境池做不到这一点。",
					"summary_ja": "SPADE は、一つの言語モデルが二つの役割を担う自己対戦型の強化学習フレームワークである。環境デザイナーは長期的な訓練環境を Gym 形式の reset と step を備えた実行可能コードとして記述し、推論エージェントはその中で学習する。狙いは、人手で整えた環境や固定された環境群では実現できない、学習者の規模に応じて広がり続ける目標分布を得ることにある。"
				},
				{
					"rank": 4,
					"title": "Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis",
					"title_zh": "训练具备化学合理性判断的大语言模型用于单步逆合成",
					"title_ja": "単一ステップ逆合成のための化学的妥当性を考慮した大規模言語モデルの訓練",
					"url": "https://huggingface.co/papers/2608.18940",
					"upvotes": 29,
					"arxiv_id": "2608.18940",
					"first_author": "Bogdan Zagribelnyy",
					"org": "Insilico Medicine",
					"summary": "Insilico Medicine trained C3LM, a chemistry model for single-step retrosynthesis, on a dataset of roughly 45.6 million verified reactions. The team proposes Top-K prompting to capture the intrinsically one-to-many nature of the problem, which single-answer benchmarks measure poorly, and pairs fine-tuning with plausibility and novelty rewards.",
					"summary_zh": "Insilico Medicine 用约 4560 万条经过验证的反应数据训练了化学模型 C3LM，用于单步逆合成。团队提出 Top-K 提示方法，以刻画该问题固有的一对多特性——这正是单一答案式评测难以衡量的部分——并将微调与合理性、新颖性奖励结合起来。",
					"summary_ja": "Insilico Medicine は、検証済みの反応およそ 4560 万件からなるデータセットで、単一ステップ逆合成向けの化学モデル C3LM を訓練した。単一の正解を前提とするベンチマークでは測りにくい一対多の性質を捉えるため Top-K プロンプティングを提案し、ファインチューニングに妥当性と新規性の報酬を組み合わせている。"
				},
				{
					"rank": 5,
					"title": "EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing",
					"title_zh": "EDITBRIDGE：面向忠实且高效的超高分辨率图像编辑",
					"title_ja": "EDITBRIDGE: 忠実かつ効率的な超高解像度画像編集に向けて",
					"url": "https://huggingface.co/papers/2608.18063",
					"upvotes": 21,
					"arxiv_id": "2608.18063",
					"first_author": "Jiayi Song",
					"summary": "EDITBRIDGE targets image editing above 1K resolution, where quadratic attention cost and memory push most diffusion models into a two-stage workaround: edit small, then upscale. The authors say that pipeline hallucinates details contradicting the high-resolution source and leaves texture over-smoothed or over-sharpened, and propose editing at full resolution instead.",
					"summary_zh": "EDITBRIDGE 面向 1K 以上分辨率的图像编辑。由于注意力计算量随分辨率平方增长且显存开销高昂，多数扩散模型只能采用两段式变通方案：先低分辨率编辑，再单独超分。作者指出这一流程会臆造出与高分辨率原图相矛盾的细节，并使纹理过度平滑或过度锐化，因此主张直接在原始分辨率上编辑。",
					"summary_ja": "EDITBRIDGE は 1K を超える解像度での画像編集を対象とする。注意機構の計算量が解像度の二乗で増え、メモリ要求も大きいため、多くの拡散モデルは低解像度で編集してから超解像するという二段構えに頼っている。著者らは、この方式が高解像度の元画像と矛盾する細部を生み、テクスチャが過度に平滑化または先鋭化されると指摘し、元の解像度のまま編集する方法を提案する。"
				}
			],
			"updated_at": "2026-08-20 13:07 PDT"
		},
		{
			"date": "2026-08-19",
			"items": [
				{
					"rank": 1,
					"title": "Demystifying Agent Skills: Why They Work-Until They Don't",
					"title_zh": "解构 Agent Skills：它们为何有效，又在何处失效",
					"title_ja": "エージェントスキルの解明: なぜ機能し、どこで機能しなくなるのか",
					"url": "https://huggingface.co/papers/2608.14036",
					"arxiv_id": "2608.14036",
					"upvotes": 102,
					"first_author": "Zhiyuan Jiang",
					"org": "University of California at San Diego",
					"summary": "A UC San Diego team ran controlled experiments to isolate when skills actually help LLM agents and when they fail. The study varies representation, outcome annotation, retrieval difficulty and cross-framework robustness across several benchmarks, agent harnesses and models. It argues aggregate task-success numbers hide where skills break down.",
					"summary_zh": "加州大学圣迭戈分校的团队通过受控实验，厘清技能（skills）在何时真正提升 LLM 智能体表现、又在何时失效。研究在多个基准、智能体框架与模型上分别考察了表示方式、结果标注、检索难度与跨框架稳健性的影响。作者认为，仅看任务成功率的汇总数字会掩盖技能失效的具体位置。",
					"summary_ja": "カリフォルニア大学サンディエゴ校のチームが、スキルがLLMエージェントを本当に助けるのはいつで、失敗するのはどこかを制御実験で切り分けた。複数のベンチマーク、エージェントハーネス、モデルにわたり、表現方法、結果アノテーション、検索の難易度、フレームワーク間の頑健性を個別に変化させて検証している。集計されたタスク成功率では破綻箇所が見えないと論じる。"
				},
				{
					"rank": 2,
					"title": "Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements",
					"title_zh": "Agentic ESOpt：以极低 GPU 需求微调长时程 LLM 智能体",
					"title_ja": "Agentic ESOpt: 最小限のGPUで長期タスクのLLMエージェントを微調整する",
					"url": "https://huggingface.co/papers/2608.17310",
					"arxiv_id": "2608.17310",
					"upvotes": 90,
					"first_author": "Zhi Zheng",
					"org": "National University of Singapore",
					"summary": "Researchers at the National University of Singapore argue for evolution strategies over reinforcement learning when fine-tuning long-horizon LLM agents. They say ES avoids the heavyweight backpropagation stack that keeps RL from scaling to larger models, and sidesteps the credit-assignment problem that branching, sparsely rewarded trajectories create.",
					"summary_zh": "新加坡国立大学的研究者主张，在微调长时程 LLM 智能体时，进化策略（ES）比强化学习更合适。他们指出，ES 无需强化学习那套依赖反向传播的重型训练栈，因而可扩展到更大的模型，同时也绕开了分支繁多、奖励稀疏的轨迹所带来的信用分配难题。",
					"summary_ja": "シンガポール国立大学の研究者らは、長期タスクを担うLLMエージェントの微調整には強化学習より進化戦略（ES）が適していると主張する。ESは逆伝播に依存する重い学習スタックを必要としないため大規模モデルにも適用でき、分岐が多く報酬の疎な軌跡で生じる信用割当の問題も回避できるとしている。"
				},
				{
					"rank": 3,
					"title": "Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation",
					"title_zh": "Embodied-Navigator：以指点、思考、记忆与对齐实现高效导航",
					"title_ja": "Embodied-Navigator: 指し示し、考え、記憶し、整合させる効率的なナビゲーション",
					"url": "https://huggingface.co/papers/2608.17512",
					"arxiv_id": "2608.17512",
					"upvotes": 43,
					"first_author": "Hongyan Feng",
					"org": "ZJU-OmniAI",
					"summary": "A ZJU-OmniAI team introduces TAMP-Nav, which recasts embodied navigation as 2D visual prompting so a vision-language model only has to select pixels. The formulation avoids the unnatural action spaces that clash with a VLM's 2D pre-training priors, and pairs it with more flexible reasoning schedules and memory management.",
					"summary_zh": "ZJU-OmniAI 团队提出 TAMP-Nav，将具身导航重构为二维视觉提示任务，使视觉语言模型只需选择像素点即可。该形式化避免了与 VLM 二维预训练先验相冲突的非自然动作空间，并配以更灵活的推理调度与记忆管理机制。",
					"summary_ja": "ZJU-OmniAIのチームはTAMP-Navを提案し、身体性ナビゲーションを2次元の視覚プロンプトとして定式化することで、視覚言語モデルはピクセルを選ぶだけでよくなる。この設計はVLMの2次元事前学習の事前知識と衝突する不自然な行動空間を避け、柔軟な推論スケジュールと記憶管理を組み合わせる。"
				},
				{
					"rank": 4,
					"title": "AVA-Encoder: Towards Agent-Native Video Representation Learning",
					"title_zh": "AVA-Encoder：面向智能体原生的视频表征学习",
					"title_ja": "AVA-Encoder: エージェントネイティブな動画表現学習に向けて",
					"url": "https://huggingface.co/papers/2608.12313",
					"arxiv_id": "2608.12313",
					"upvotes": 35,
					"first_author": "Chuyue Li",
					"org": "Qwen Business Unit",
					"summary": "The Qwen Business Unit proposes AVA-Encoder, which converts a video into a knowledge-graph representation and then reconstructs the video from it. The agentic auto-encoding setup aims to give creative agents a structure they can reason over and manipulate, rather than an opaque latent, so they can learn from high-quality human films.",
					"summary_zh": "Qwen 业务部门提出 AVA-Encoder，先把视频转换为知识图谱表示，再由该表示重建出视频。这种“智能体式自编码”旨在为创作类智能体提供可推理、可操作的结构，而非不透明的隐空间表示，从而让它们能够从高质量的人类影片中学习。",
					"summary_ja": "Qwenビジネスユニットは、動画をナレッジグラフ表現に変換し、そこから動画を再構成するAVA-Encoderを提案した。このエージェント的オートエンコーディングは、不透明な潜在表現ではなく推論・操作が可能な構造を創作エージェントに与え、質の高い人間の映像作品から学べるようにすることを狙う。"
				},
				{
					"rank": 5,
					"title": "An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models",
					"title_zh": "像素空间文生图扩散模型训练的实证研究",
					"title_ja": "ピクセル空間でのテキスト画像生成拡散モデル学習に関する実証研究",
					"url": "https://huggingface.co/papers/2608.16887",
					"arxiv_id": "2608.16887",
					"upvotes": 30,
					"first_author": "Dengyang Jiang",
					"org": "Tongyi-MAI",
					"summary": "Tongyi-MAI reports a large-scale study of text-to-image diffusion trained directly in pixel space rather than in a latent space. The authors find that direct pixel-space pre-training converges substantially more slowly at scale, and propose a latent-to-pixel strategy that acquires the representation in latent space first.",
					"summary_zh": "Tongyi-MAI 发布了一项关于直接在像素空间（而非隐空间）训练文生图扩散模型的大规模实证研究。作者发现，大规模直接在像素空间预训练的收敛速度明显更慢，并据此提出先在隐空间获得表征、再迁移到像素空间的“隐空间到像素”策略。",
					"summary_ja": "Tongyi-MAIは、潜在空間ではなくピクセル空間で直接学習するテキスト画像生成拡散モデルの大規模な実証研究を報告した。著者らは、大規模な事前学習をピクセル空間で直接行うと収束が大幅に遅くなることを見いだし、まず潜在空間で表現を獲得する「潜在からピクセルへ」の戦略を提案している。"
				}
			],
			"updated_at": "2026-08-19 13:05 PDT"
		},
		{
			"date": "2026-08-18",
			"items": [
				{
					"rank": 1,
					"title": "HarnessEval-W: Agentifying the Evaluation of Visual Worlds",
					"title_zh": "HarnessEval-W：将视觉世界的评测智能体化",
					"title_ja": "HarnessEval-W: 視覚世界の評価のエージェント化",
					"url": "https://huggingface.co/papers/2608.16859",
					"arxiv_id": "2608.16859",
					"upvotes": 109,
					"first_author": "Weiliang Chen",
					"org": "MirroS",
					"summary": "HarnessEval-W is an evaluation pipeline that judges world model rollouts with an agent harness instead of brute-force metrics. It produces a reasoning chain justifying each score, so violations of physics, causality and world state can be examined and verified rather than reduced to a scalar. It is the day's most upvoted paper at 109 votes.",
					"summary_zh": "HarnessEval-W 是一套评测流程，用智能体 harness 而非暴力计算的指标来判定世界模型的推演结果。它为每个分数生成可供检验的推理链，使物理、因果与世界状态上的破绽能够被审查和验证，而不是被压缩成一个标量分数。该论文以 109 票成为当日票数最高的一篇。",
					"summary_ja": "HarnessEval-W は、ワールドモデルのロールアウトを総当たり的な指標ではなくエージェントのハーネスで評価するパイプラインである。各スコアの根拠となる推論の連鎖を出力するため、物理・因果・世界状態の破綻をスカラー値に還元せず検証できる。本日最多となる109票を集めた。"
				},
				{
					"rank": 2,
					"title": "StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling",
					"title_zh": "StateM：通过 harness 扩展在 Terminal-Bench 2.1 上达到 95.3% 原始准确率，或 15 美元的前沿运行",
					"title_ja": "StateM: ハーネススケーリングによるTerminal-Bench 2.1での95.3%の生正答率、あるいは15ドルのフロンティア実行",
					"url": "https://huggingface.co/papers/2608.15089",
					"arxiv_id": "2608.15089",
					"upvotes": 51,
					"first_author": "Ziheng Qin",
					"summary": "StateM is an agent-native runtime that organizes long-horizon execution around durable states, phase-local context, checked transitions and recoverable runbooks, leaving model weights untouched. The authors report 95.3 percent raw accuracy on Terminal-Bench 2.1 and a frontier run costing 15 dollars, framing it as a case for scaling the harness rather than the model.",
					"summary_zh": "StateM 是一个面向智能体的运行时，围绕持久状态、阶段局部上下文、受检的状态转移和可恢复的操作手册来组织长程执行，且不改动模型权重。作者报告其在 Terminal-Bench 2.1 上取得 95.3% 的原始准确率，一次前沿运行成本为 15 美元，并以此论证应当扩展 harness 而非扩展模型。",
					"summary_ja": "StateM は、永続的な状態、フェーズ局所のコンテキスト、検査付きの遷移、復旧可能なランブックを軸に長期タスクの実行を組み立てるエージェント向けランタイムで、モデルの重みは変更しない。著者らは Terminal-Bench 2.1 で95.3%の生の正答率と15ドルのフロンティア実行を報告し、モデルではなくハーネスを拡張する立場を示している。"
				},
				{
					"rank": 3,
					"title": "Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search",
					"title_zh": "大型发现模型：基于实证的模型驱动开放式搜索",
					"title_ja": "大規模ディスカバリーモデル: 実証に基づくモデルベースの開放的探索",
					"url": "https://huggingface.co/papers/2608.15669",
					"arxiv_id": "2608.15669",
					"upvotes": 48,
					"first_author": "Zhongwei Yu",
					"org": "yangtze-ailab",
					"summary": "The Large Discovery Model is a recurrent architecture for open-ended search over structured hypothesis spaces such as molecules, protein sequences and programs, where each candidate is expensive to evaluate. The authors argue that language model likelihoods and self-assessments are unreliable proxies for the objective and for calibrated uncertainty, and ground the search empirically instead.",
					"summary_zh": "Large Discovery Model 是一种循环架构，用于在分子、蛋白质序列和程序等结构化且开放的假设空间中搜索，这类空间中每个候选的评估成本都很高。作者认为语言模型的似然与自评并不能可靠代表目标函数或校准后的不确定性，因而改用实证数据来支撑搜索。",
					"summary_ja": "Large Discovery Model は、分子・タンパク質配列・プログラムといった構造化された開放的な仮説空間を探索する再帰型アーキテクチャで、候補ごとの評価コストが高い設定を対象とする。著者らは、言語モデルの尤度や自己評価は目的関数や較正された不確実性の代理として信頼できないとし、実証データに基づく探索を提案する。"
				},
				{
					"rank": 4,
					"title": "Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization",
					"title_zh": "学尚未掌握的，而非已经掌握的：面向多奖励策略优化的饱和感知优势重加权",
					"title_ja": "習得済みではなく残されたものを学ぶ: 多目的報酬の方策最適化に向けた飽和認識型アドバンテージ再重み付け",
					"url": "https://huggingface.co/papers/2608.16072",
					"arxiv_id": "2608.16072",
					"upvotes": 43,
					"first_author": "Yixuan Wang",
					"org": "University of California at San Diego",
					"summary": "The paper targets multi-reward reinforcement learning for post-training reasoners, where reward vectors are usually collapsed into a fixed weighted sum before group-wise standardization. The authors show this hands identical advantages to rollouts with different reward profiles and keeps spending gradient on objectives that are already saturated.",
					"summary_zh": "该论文针对推理模型后训练中的多奖励强化学习：现有做法通常先把奖励向量按固定权重求和，再做组内标准化。作者指出，这会让奖励构成完全不同的采样获得相同的优势值，并持续把梯度花在已经饱和的目标上。",
					"summary_ja": "本論文は、推論モデルの事後学習における多目的報酬の強化学習を扱う。従来手法は報酬ベクトルを固定重みで加算してからグループ内で標準化するため、報酬の内訳が異なるロールアウトに同一のアドバンテージが与えられ、すでに飽和した目的にも勾配が費やされ続けると著者らは指摘する。"
				},
				{
					"rank": 5,
					"title": "MOSS-VL Technical Report",
					"title_zh": "MOSS-VL 技术报告",
					"title_ja": "MOSS-VL テクニカルレポート",
					"url": "https://huggingface.co/papers/2608.15045",
					"arxiv_id": "2608.15045",
					"upvotes": 38,
					"first_author": "Pengyu Wang",
					"org": "OpenMOSS",
					"summary": "OpenMOSS released MOSS-VL, an open vision-language model family designed to perceive while it speaks. The language decoder attends to vision only through gated cross-attention, so incoming frames can be handled during generation, and a synthesized interaction corpus supervises when to speak, stay silent or revise. Real-time training sits in one light final stage over an offline foundation.",
					"summary_zh": "OpenMOSS 发布了开源视觉语言模型系列 MOSS-VL，把边看边说的实时交互作为一等能力来设计。语言解码器仅通过门控交叉注意力接入视觉，因而可以在生成过程中处理新到的画面；合成的交互语料则监督模型何时开口、何时沉默、何时修正。实时相关的训练集中在离线基座之上的最后一个轻量阶段。",
					"summary_ja": "OpenMOSS は、話しながら見るリアルタイム対話を第一級の機能として設計したオープンな視覚言語モデル群 MOSS-VL を公開した。言語デコーダはゲート付きクロスアテンションのみで視覚を参照するため、生成中に届くフレームを処理できる。合成した対話コーパスが、話す・黙る・言い直すタイミングを教師信号として与える。"
				}
			],
			"updated_at": "2026-08-18 13:07 PDT"
		},
		{
			"date": "2026-08-17",
			"items": [
				{
					"rank": 1,
					"title": "Self-Supervised Visual On-Policy Distillation",
					"url": "https://huggingface.co/papers/2608.14144",
					"arxiv_id": "2608.14144",
					"upvotes": 144,
					"first_author": "Yijiang Li",
					"org": "University of California at San Diego",
					"title_zh": "自监督的视觉在线策略蒸馏",
					"title_ja": "自己教師あり視覚オンポリシー蒸留",
					"summary": "The paper asks where the teacher-student asymmetry in visual on-policy distillation can come from when no privileged supervision is available. Instead of handing the teacher extra information such as reference answers or ground-truth regions of interest, the authors subtract information from the student, which they report yields the same effective learning signal for free.",
					"summary_zh": "论文提出的问题是：当没有任何特权监督信号可用时，视觉在线策略蒸馏所依赖的师生信息不对称从何而来。作者不再向教师模型追加参考答案或真值感兴趣区域等额外信息，而是反过来削减学生模型可获得的信息，并称这样能够免费获得同等有效的学习信号。",
					"summary_ja": "本論文は、特権的な教師信号が一切ない状況で、視覚オンポリシー蒸留に必要な教師と生徒の非対称性をどこから得るかを問う。著者らは参照解答や正解の注目領域といった追加情報を教師に与えるのではなく、逆に生徒から情報を差し引くことで、同等の学習信号をコストなしに得られると報告している。"
				},
				{
					"rank": 2,
					"title": "Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence",
					"url": "https://huggingface.co/papers/2608.11341",
					"arxiv_id": "2608.11341",
					"upvotes": 22,
					"first_author": "Brian Wang",
					"org": "Apodex",
					"title_zh": "Apodex Discovery：面向发现型人工智能的评估与构建的现实基准与环境",
					"title_ja": "Apodex Discovery: 発見型人工知能を評価・構築するための現実ベンチマークと環境",
					"summary": "Apodex Discovery is a framework of benchmarks and environments for AI aimed at open-ended discovery rather than pre-specified tasks, built around what the 29 authors call a heavy-duty solver. Their argument is that frontier models already handle hard problems once the goal, tools and success criteria are executable, and the missing step is turning real-world challenges into that form.",
					"summary_zh": "Apodex Discovery 是一套面向开放式发现（而非预先设定任务）的 AI 基准与环境框架，核心是 29 位作者所称的 heavy-duty solver。他们的论点是：一旦目标、工具与成功标准可被执行和验证，前沿模型已能解决困难问题，真正缺失的一步是把现实世界的挑战转化为这种可执行的形式。",
					"summary_ja": "Apodex Discovery は、あらかじめ規定されたタスクではなく開かれた発見を対象とする AI のためのベンチマークと環境の枠組みで、29 名の著者が heavy-duty solver と呼ぶ仕組みを中核に据える。目標・ツール・成功基準が実行可能な形になれば最先端モデルは難問を解けており、欠けているのは現実の課題をその形に変換する工程だ、というのが主張である。"
				},
				{
					"rank": 3,
					"title": "Marionette: Predicting World States, Rendering Geometry, Painting Appearance",
					"url": "https://huggingface.co/papers/2608.14530",
					"arxiv_id": "2608.14530",
					"upvotes": 22,
					"first_author": "Zian Meng",
					"title_zh": "Marionette：预测世界状态、渲染几何、绘制外观",
					"title_ja": "Marionette: 世界状態の予測、ジオメトリのレンダリング、外観の描画",
					"summary": "Marionette splits an interactive game world model into three jobs: predicting the evolving world state, computing geometry with a fixed zero-parameter renderer, and leaving the neural network to synthesize appearance. The design targets long-horizon drift, where models that autoregress pixels or latents accumulate errors in pose, geometry and occlusion until control breaks down.",
					"summary_zh": "Marionette 把交互式游戏世界模型拆成三件事：预测不断演化的世界状态、由固定的零参数渲染器计算几何、再交给神经网络合成外观。该设计针对的是长时程漂移问题：直接对像素或隐变量做自回归的模型会不断累积姿态、几何和遮挡上的误差，最终导致可控性崩坏。",
					"summary_ja": "Marionette は、インタラクティブなゲーム世界モデルを三つの役割に分解する。変化する世界状態の予測、パラメータを持たない固定レンダラーによるジオメトリ計算、そしてニューラルネットワークによる外観の合成である。狙いは長期的なドリフトの抑制で、ピクセルや潜在表現を自己回帰する方式では姿勢・幾何・遮蔽の誤差が蓄積し、制御が破綻していく。"
				},
				{
					"rank": 4,
					"title": "DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data",
					"url": "https://huggingface.co/papers/2608.13517",
					"arxiv_id": "2608.13517",
					"upvotes": 21,
					"first_author": "Peter Schneider-Kamp",
					"org": "University of Southern Denmark (SDU)",
					"title_zh": "DFM Mimir v1：仅用可授权后训练数据、在 10 亿参数上实现前沿性能的开放 HRM",
					"title_ja": "DFM Mimir v1: 許諾されたポストトレーニングデータのみで10億パラメータのフロンティア性能を実現するオープンHRM",
					"summary": "Mimir v1 is a 1-billion-parameter model built on the Hierarchical Reasoning Model architecture and trained from scratch on a mixture of 161 permissibly licensed datasets. It outperforms the original HRM-Text 1B and sets a new state of the art for Danish, which the authors offer as evidence that ethically sourced data need not forfeit competitive results at this scale.",
					"summary_zh": "Mimir v1 是一个基于分层推理模型（HRM）架构的 10 亿参数模型，完全从零开始训练，使用了 161 个具备合规授权的数据集组成的混合语料。它的表现优于原始的 HRM-Text 1B，并在丹麦语上刷新了最好成绩；作者以此说明，在这一参数规模上，坚持合规取得的数据并不意味着放弃有竞争力的效果。",
					"summary_ja": "Mimir v1 は階層的推論モデル（HRM）アーキテクチャに基づく 10 億パラメータのモデルで、利用許諾の明確な 161 のデータセットを混合したコーパスでゼロから学習された。オリジナルの HRM-Text 1B を上回り、デンマーク語では新たな最高性能を記録しており、著者らはこの規模であれば倫理的に調達したデータでも競争力を犠牲にしないことの証拠だとしている。"
				},
				{
					"rank": 5,
					"title": "CPI-Bench: A Comprehensive, Practical and Intelligent Benchmark for Real-World Image Editing",
					"url": "https://huggingface.co/papers/2608.14546",
					"arxiv_id": "2608.14546",
					"upvotes": 12,
					"first_author": "Qinye Zhou",
					"org": "Alibaba",
					"title_zh": "CPI-Bench：面向真实场景图像编辑的综合、实用与智能基准",
					"title_ja": "CPI-Bench: 実世界の画像編集に向けた包括的・実用的で知的なベンチマーク",
					"summary": "CPI-Bench is an image editing benchmark built for deployment-like conditions, covering multi-image edits, demanding reasoning instructions and practical settings. The 20 authors argue that existing benchmarks stay confined to simple single-image tasks, which leaves them unable to separate the performance of today's editing models.",
					"summary_zh": "CPI-Bench 是一套面向接近实际部署条件的图像编辑基准，覆盖多图编辑、对推理要求较高的指令以及实用场景。20 位作者认为，现有基准仍局限于简单的单图任务，因而无法有效区分当下各类图像编辑模型的表现差异。",
					"summary_ja": "CPI-Bench は実運用に近い条件を想定した画像編集ベンチマークで、複数画像の編集、高度な推論を要する指示、実践的な利用場面を対象とする。20 名の著者は、既存のベンチマークが単純な単一画像タスクにとどまっており、現在の編集モデル同士の性能差を判別できていないと指摘する。"
				}
			],
			"updated_at": "2026-08-17 13:06 PDT"
		},
		{
			"date": "2026-08-15",
			"items": [
				{
					"rank": 1,
					"title": "LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time",
					"title_zh": "LiveAnimate：实时稳定的长时流式人体动画生成",
					"title_ja": "LiveAnimate: 実時間で安定した長尺ストリーミング人物アニメーション",
					"url": "https://huggingface.co/papers/2608.11745",
					"summary": "LiveAnimate turns a single reference image and a driving pose stream into human animation in real time, aimed at live streaming, telepresence and virtual avatars where diffusion systems need minutes per clip. It is built on a 14B-parameter video diffusion transformer with two-stage training, which the authors call the first system to pair real-time streaming with stable long-form generation at that scale.",
					"summary_zh": "LiveAnimate 可将单张参考图像与驱动姿态流实时转换为人体动画，面向直播、远程呈现和虚拟形象等场景——现有扩散方法生成一段片段往往需要数分钟。该系统基于 140 亿参数的视频扩散 Transformer，采用两阶段训练，作者称其是首个在该规模上兼顾实时流式生成与长时稳定输出的系统。",
					"summary_ja": "LiveAnimate は、1 枚の参照画像と駆動ポーズ列から人物アニメーションをリアルタイムに生成し、ライブ配信やテレプレゼンス、バーチャルアバターを想定する。従来の拡散モデルは 1 クリップの生成に数分を要していた。140 億パラメータの動画拡散 Transformer を二段階学習で構築し、この規模でリアルタイム配信と長尺の安定生成を両立した初のシステムだと著者は述べている。",
					"upvotes": 12,
					"arxiv_id": "2608.11745",
					"first_author": "Yuxuan Zhang",
					"org": "Qwen Business Unit"
				},
				{
					"rank": 2,
					"title": "Full-bandwidth transformer",
					"title_zh": "全带宽 Transformer",
					"title_ja": "フルバンド幅 Transformer",
					"url": "https://huggingface.co/papers/2608.08888",
					"summary": "Microsoft Research proposes widening the vertical feedback channel in autoregressive transformers, where today only the sampled token returns to the bottom of the stack and the top-layer hidden state is discarded. The full-bandwidth transformer fuses that previous hidden state with the sampled token embedding through a gated linear unit at each decoding step.",
					"summary_zh": "微软研究院提出拓宽自回归 Transformer 中的纵向反馈通道：现有解码过程只把采样得到的 token 送回堆栈底部，而顶层隐状态被直接丢弃。全带宽 Transformer 在每个解码步通过门控线性单元，将上一步的顶层隐状态与采样 token 的嵌入融合。",
					"summary_ja": "Microsoft Research は、自己回帰 Transformer の垂直方向のフィードバック経路を広げる手法を提案した。現在の復号では、サンプリングされたトークンだけがスタック下部に戻り、最上層の隠れ状態は破棄されている。フルバンド幅 Transformer は各復号ステップで、直前の最上層隠れ状態とトークン埋め込みをゲート付き線形ユニットで融合する。",
					"upvotes": 10,
					"arxiv_id": "2608.08888",
					"first_author": "Xi Wang",
					"org": "Microsoft Research"
				},
				{
					"rank": 3,
					"title": "An AI4AI Framework for Visual Token Pruning",
					"title_zh": "面向视觉 token 剪枝的 AI4AI 框架",
					"title_ja": "視覚トークン枝刈りのための AI4AI フレームワーク",
					"url": "https://huggingface.co/papers/2608.07193",
					"summary": "Visual token pruning cuts the inference cost of multimodal LLMs, but the methods in use rely on hand-tuned heuristics and expert trial and error. This paper asks whether large language models can design the pruning algorithms themselves, and proposes an AI4AI framework that searches the design space as pruning objectives, budgets and architectures multiply.",
					"summary_zh": "视觉 token 剪枝可以显著降低多模态大模型的推理成本，但现有方法大多依赖手工设计的启发式规则和专家反复试错。这篇论文提出让大语言模型自行设计剪枝算法，构建了一个 AI4AI 框架，在剪枝目标、预算和模型架构不断增多的设计空间中自动搜索。",
					"summary_ja": "視覚トークンの枝刈りはマルチモーダル大規模言語モデルの推論コストを大きく下げられるが、既存手法は手作業のヒューリスティクスと専門家の試行錯誤に依存している。本論文は大規模言語モデル自身に枝刈りアルゴリズムを設計させられるかを問い、目的・予算・アーキテクチャが多様化する設計空間を探索する AI4AI フレームワークを提案する。",
					"upvotes": 7,
					"arxiv_id": "2608.07193",
					"first_author": "Zhen Liu"
				},
				{
					"rank": 4,
					"title": "H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models",
					"title_zh": "H2R-Bench：世界模型中人到机器人操作视频生成的基准",
					"title_ja": "H2R-Bench: 世界モデルにおける人間からロボットへの操作動画生成ベンチマーク",
					"url": "https://huggingface.co/papers/2608.13049",
					"summary": "Robot demonstration data is expensive to collect, while egocentric human manipulation video is abundant but hard to transfer across embodiments. H2R-Bench, from Shanghai Jiao Tong University, evaluates how well video world models synthesize robot-centric manipulation clips from human footage, a cross-embodiment capability the authors say remains largely untested.",
					"summary_zh": "机器人示教数据采集成本高昂，而第一人称的人类操作视频资源丰富，却难以跨形态迁移到机器人。上海交通大学提出的 H2R-Bench 用于评估视频世界模型能否从人类视频合成以机器人为主体的操作片段，作者指出这一跨形态能力此前基本未被系统检验。",
					"summary_ja": "ロボットの実演データ収集は高コストである一方、一人称視点の人間の操作動画は豊富だが、身体構造の違いから転用が難しい。上海交通大学の H2R-Bench は、動画ワールドモデルが人間の映像からロボット視点の操作クリップをどれだけ合成できるかを評価する。著者はこの身体間転移能力がほとんど検証されてこなかったと指摘する。",
					"upvotes": 7,
					"arxiv_id": "2608.13049",
					"first_author": "Dingyi Rong",
					"org": "Shanghai Jiao Tong University"
				},
				{
					"rank": 5,
					"title": "Thought-Level Beam Search for Reasoning",
					"title_zh": "面向推理的思维级束搜索",
					"title_ja": "推論のための思考レベルビームサーチ",
					"url": "https://huggingface.co/papers/2608.08020",
					"summary": "Princeton researchers frame test-time reasoning as a compute allocation problem over partial trajectories rather than a question of how much compute to spend. They argue that parallel sampling treats traces independently and creates memory bottlenecks, while subtractive pruning starves promising branches, and propose beam search over whole thoughts instead of tokens.",
					"summary_zh": "普林斯顿大学的研究者把测试时推理重新表述为对部分推理轨迹的算力分配问题，而不是「该花多少算力」的问题。他们指出并行采样把各条轨迹当作彼此独立，会造成严重的显存瓶颈，而剪枝式方法又会「饿死」有希望的分支，为此提出以完整思维而非 token 为单位的束搜索。",
					"summary_ja": "プリンストン大学の研究者らは、テスト時推論を「どれだけ計算資源を使うか」ではなく、途中までの推論軌跡にどう配分するかという問題として定式化した。並列サンプリングは各軌跡を独立に扱うためメモリのボトルネックを生み、枝刈り型の手法は有望な分岐を早期に切り捨てると指摘し、トークンではなく思考単位でのビームサーチを提案する。",
					"upvotes": 6,
					"arxiv_id": "2608.08020",
					"first_author": "Lijie Yang",
					"org": "Princeton University"
				}
			],
			"updated_at": "2026-08-15 13:05 PDT"
		},
		{
			"date": "2026-08-14",
			"items": [
				{
					"rank": 1,
					"title": "Alaya-EVOKE: From Linear-Scaling Supervision to Endless World",
					"url": "https://huggingface.co/papers/2608.13546",
					"arxiv_id": "2608.13546",
					"upvotes": 81,
					"first_author": "Yuanyang Yin",
					"title_zh": "Alaya-EVOKE：从线性扩展的监督到无尽世界",
					"title_ja": "Alaya-EVOKE: 線形スケーリングの教師信号から終わりなき世界へ",
					"summary": "Evoke is an interactive world model that externalizes persistent world state instead of holding history in the denoiser context or key-value cache, whose cost grows with session length. Scene geometry is maintained outside the model and the teacher is redesigned for long-horizon interactive generation. The paper drew 81 upvotes, the most of the day.",
					"summary_zh": "Evoke 是一种交互式世界模型，它将持久化的世界状态外置，而不是把历史保存在去噪器上下文或键值缓存中，后者的开销会随会话长度增长。场景几何在模型之外维护，同时重新设计教师模型以支持长时程交互式生成。该论文获得 81 个赞，为当日最高。",
					"summary_ja": "Evoke は、履歴をデノイザーのコンテキストや KV キャッシュに保持する代わりに、永続的な世界状態を外部化する対話型ワールドモデルである。シーンの幾何情報はモデルの外部で保持され、教師モデルは長時程の対話生成向けに再設計された。同論文は当日最多の 81 票を集めた。"
				},
				{
					"rank": 2,
					"title": "DarwinX: Evolving Agent Harnesses Through Natural Selection",
					"url": "https://huggingface.co/papers/2608.07545",
					"arxiv_id": "2608.07545",
					"upvotes": 54,
					"first_author": "Yifan Zhang",
					"org": "Salesforce AI Research",
					"title_zh": "DarwinX：通过自然选择进化智能体框架",
					"title_ja": "DarwinX: 自然選択によるエージェント・ハーネスの進化",
					"summary": "DarwinX treats agent self-improvement as selection over a population of harnesses - prompts, tools, skills and control flow - with the model weights frozen. A preserve-and-extend contract admits only variants that extend coverage without regressing other tasks, while an archive keeps alternative lineages available for recombination.",
					"summary_zh": "DarwinX 将智能体的自我改进视为对一组框架（提示词、工具、技能与控制流）种群的选择过程，模型权重保持冻结。其“保留并扩展”约定只接纳能扩大覆盖范围且不导致其他任务退化的变体，同时用存档保留其他谱系以供重组。",
					"summary_ja": "DarwinX は、モデルの重みを凍結したまま、プロンプト・ツール・スキル・制御フローからなるハーネスの集団に対する選択としてエージェントの自己改善を捉える。「保存と拡張」の契約により、他タスクを劣化させずにカバー範囲を広げる変異のみを採用し、アーカイブが別系統を組み換え用に保持する。"
				},
				{
					"rank": 3,
					"title": "How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review",
					"url": "https://huggingface.co/papers/2608.08975",
					"arxiv_id": "2608.08975",
					"upvotes": 38,
					"first_author": "Ming Li",
					"org": "University of Maryland",
					"title_zh": "修辞如何对 AI 审稿人实施奖励攻击？剖析 AI 同行评审中的修辞敏感性",
					"title_ja": "レトリックはAIレビュアーをどう報酬ハックするか: AI査読における修辞的感受性の解剖",
					"summary": "The authors rewrote 120 anonymized ICLR 2026 submissions along six rhetorical dimensions while preserving the reported scientific content, producing a controlled corpus of 4,200 manuscripts. Five LLM reviewers then scored the results under standard and strict protocols, measuring how far presentation alone shifts AI review judgments.",
					"summary_zh": "作者在保持所报告科学内容不变的前提下，沿六个修辞维度改写了 120 篇匿名的 ICLR 2026 投稿，构建出包含 4200 篇稿件的受控语料。随后由五个大模型审稿人在标准与严格两种协议下评分，以衡量仅凭表达方式就能在多大程度上改变 AI 的评审判断。",
					"summary_ja": "著者らは、報告された科学的内容を保ったまま、匿名化された ICLR 2026 投稿 120 本を 6 つの修辞的次元に沿って書き換え、4200 本の統制コーパスを構築した。5 つの LLM レビュアーが標準および厳格なプロトコルで採点し、提示の仕方だけで AI の査読判断がどこまで動くかを測定した。"
				},
				{
					"rank": 4,
					"title": "Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence",
					"url": "https://huggingface.co/papers/2608.12743",
					"arxiv_id": "2608.12743",
					"upvotes": 27,
					"first_author": "Haokai Zhang",
					"org": "Zhejiang University",
					"title_zh": "空间记忆智能体：面向空间智能的经验驱动流程记忆",
					"title_ja": "空間記憶エージェント: 空間知能のための経験に基づく手順記憶",
					"summary": "The paper asks whether a frozen vision-language model can improve its spatial reasoning through accumulated experience alone. It proposes procedure memory grounded in past attempts as a third route, complementary to post-training methods such as fine-tuning and to agents that call external depth and 3D reconstruction tools.",
					"summary_zh": "该论文探讨冻结参数的视觉语言模型能否仅凭经验积累提升空间推理能力。作者提出以过往尝试为依据的流程记忆，作为第三条路径，与微调等后训练方法以及调用深度估计、三维重建等外部工具的智能体形成互补。",
					"summary_ja": "本論文は、重みを凍結した視覚言語モデルが経験の蓄積だけで空間推論を改善できるかを問う。過去の試行に基づく手順記憶を第三の道として提案し、ファインチューニングなどの事後学習や、深度推定・3D 復元といった外部ツールを呼ぶエージェント手法を補完する。"
				},
				{
					"rank": 5,
					"title": "Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus",
					"url": "https://huggingface.co/papers/2608.12149",
					"arxiv_id": "2608.12149",
					"upvotes": 16,
					"first_author": "Zunhai Su",
					"org": "Startlux",
					"title_zh": "混合线性注意力大模型中的巨量激活：注意力前尖峰与尖峰间平台",
					"title_ja": "ハイブリッド線形アテンションLLMにおける巨大活性化: アテンション前スパイクとスパイク間プラトー",
					"summary": "This is the first systematic study of massive activations in layer-interleaved hybrid linear attention models, and it identifies two recurring shapes: spikes that appear immediately before full attention layers, and plateaus that persist across the intervening linear attention layers. The organization recurred across five linear attention architectures.",
					"summary_zh": "这是首个针对层间交错式混合线性注意力模型中巨量激活的系统性研究，识别出两种反复出现的形态：紧接在全注意力层之前出现的尖峰，以及跨越中间线性注意力层持续存在的平台。该组织形式在五种线性注意力架构中反复出现。",
					"summary_ja": "層が交互に配置されたハイブリッド線形アテンションモデルにおける巨大活性化を初めて体系的に調べた研究で、フルアテンション層の直前に現れるスパイクと、その間の線形アテンション層をまたいで持続するプラトーという二つの形態を特定した。この構造は 5 つの線形アテンション・アーキテクチャで共通して確認された。"
				}
			],
			"updated_at": "2026-08-14 13:05 PDT"
		},
		{
			"date": "2026-08-13",
			"items": [
				{
					"rank": 1,
					"title": "Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill",
					"title_zh": "Spark-to-Paper：以可组合技能实现端到端论文生成",
					"title_ja": "Spark-to-Paper: 合成可能なスキルとしての端から端までの論文生成",
					"url": "https://huggingface.co/papers/2608.11924",
					"arxiv_id": "2608.11924",
					"upvotes": 176,
					"first_author": "Zhuoyang Qian",
					"summary": "Spark-to-Paper turns a research idea into a complete paper through thirteen composable skills running inside an existing coding assistant, with no separate agent platform or orchestration service. The system retrieves literature, designs and runs experiments, revises claims against evidence and produces publication-ready figures, separating model judgment from deterministic operations.",
					"summary_zh": "Spark-to-Paper 把一个研究想法推进为完整论文，方式是在现有编程助手内部运行十三项可组合的技能，无需另建智能体平台或编排服务。系统会检索文献、设计并执行实验、依据证据修订论断，并生成可直接发表的图表，同时把模型判断与确定性操作分开处理。",
					"summary_ja": "Spark-to-Paper は、既存のコーディングアシスタント内で動く 13 の合成可能なスキルによって、研究アイデアを完成した論文へと仕上げる。専用のエージェント基盤やオーケストレーションサービスを必要とせず、文献検索、実験の設計と実行、証拠に基づく主張の修正、投稿可能な図の作成までを担い、モデルによる判断と決定的な処理を分離している。"
				},
				{
					"rank": 2,
					"title": "Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence",
					"title_zh": "Mechanist：把 AI 当作科学仪器来发现智能的内在机制",
					"title_ja": "Mechanist: 知能のメカニズムを発見する科学的計測器としての AI",
					"url": "https://huggingface.co/papers/2608.12036",
					"arxiv_id": "2608.12036",
					"upvotes": 73,
					"first_author": "Mengru Wang",
					"org": "ZJUNLP",
					"summary": "Mechanist is an agentic system that uses AI itself as a scientific instrument to autonomously discover the mechanisms behind model capabilities. The 19-author team argues that mechanistic exploration remains largely manual even as AI development accelerates and automates, widening the gap between what models can do and what researchers can understand or control.",
					"summary_zh": "Mechanist 是一套智能体系统，把 AI 本身当作科学仪器，用以自主发现模型能力背后的内在机制。这篇由 19 位作者合著的论文指出，在 AI 研发不断加速并走向自动化的同时，机制层面的探索仍主要依赖人工，使模型能力与研究者的理解、控制之间的差距进一步拉大。",
					"summary_ja": "Mechanist は、AI 自体を科学的な計測器として用い、モデルの能力を支えるメカニズムを自律的に発見するエージェント型システムである。19 名の著者は、AI 開発が加速し自動化される一方でメカニズムの探究は依然として手作業に頼っており、モデルにできることと研究者が理解・制御できることの差が広がっていると論じる。"
				},
				{
					"rank": 3,
					"title": "SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries",
					"title_zh": "SkillZip：面向可扩展智能体技能库的契约保持式图压缩",
					"title_ja": "SkillZip: 拡張可能なエージェントスキルライブラリのための契約保持型グラフ圧縮",
					"url": "https://huggingface.co/papers/2608.05604",
					"arxiv_id": "2608.05604",
					"upvotes": 72,
					"first_author": "Xingyu Tan",
					"summary": "SkillZip compresses the reusable skill packages that LLM agents load at inference time, so the smallest sufficient executable context fits a limited budget. The authors trace current failures to a unit mismatch, where skills are retrieved as whole packages but reused below that level, and compress as a graph that preserves procedural contracts and keeps routines executable.",
					"summary_zh": "SkillZip 针对大模型智能体在推理时加载的可复用技能包进行压缩，让「足够执行的最小上下文」能装进有限预算。作者认为现有系统的症结在于粒度错配：技能以整包检索，复用却发生在包以下的层级；为此他们以图的方式压缩，保留过程性契约并确保例程仍可执行。",
					"summary_ja": "SkillZip は、LLM エージェントが推論時に読み込む再利用可能なスキルパッケージを圧縮し、実行に十分な最小限のコンテキストを限られた予算に収める。著者らは、スキルがパッケージ単位で検索される一方で再利用はそれより細かい粒度で起きるという単位の不一致を問題とみなし、手続き上の契約を保ったまま実行可能性を維持するグラフ圧縮を提案する。"
				},
				{
					"rank": 4,
					"title": "Articulated Object Reconstruction from Rest-State Observation",
					"title_zh": "仅凭静止状态观测重建可动物体",
					"title_ja": "静止状態の観測からの可動物体の再構成",
					"url": "https://huggingface.co/papers/2607.27749",
					"arxiv_id": "2607.27749",
					"upvotes": 42,
					"first_author": "Daeun Lee",
					"org": "Seoul National University",
					"summary": "The paper reconstructs articulated objects from a single closed configuration, recovering both 3D geometry and the kinematic structure that governs how parts move. Existing methods need observable motion across multiple articulation states, so the Seoul National University team leans on geometry, semantics and motion priors, using an explicit mesh for cross-model verification.",
					"summary_zh": "该论文只用一个闭合状态的观测就重建出可动物体，同时恢复三维几何与决定各部件如何运动的运动学结构。既有方法需要跨多个关节状态的可观测运动，而首尔大学团队改以几何、语义与运动先验加以弥补，并用显式网格作为中间表示进行跨模型验证。",
					"summary_ja": "本論文は、閉じた 1 つの状態の観測だけから可動物体を再構成し、3D 形状と各部位の動きを規定する運動学構造の両方を復元する。既存手法は複数の関節状態にわたる観測可能な動きを必要とするため、ソウル大学のチームは幾何・意味・動きの事前知識で補い、明示的なメッシュを中間表現としてモデル間の検証に用いる。"
				},
				{
					"rank": 5,
					"title": "AdvFD: Boosting Visual Generation via Adversarial Frechet Distance Loss",
					"title_zh": "AdvFD：用对抗式 Frechet 距离损失提升视觉生成质量",
					"title_ja": "AdvFD: 敵対的フレシェ距離損失による視覚生成の強化",
					"url": "https://huggingface.co/papers/2608.11205",
					"arxiv_id": "2608.11205",
					"upvotes": 24,
					"first_author": "Mingju Gao",
					"org": "Kolors Team, Kuaishou Technology",
					"summary": "AdvFD attacks Frechet hacking, where a generator keeps improving its target Frechet score while visual quality and alignment in other feature spaces stagnate or degrade. The Kuaishou Kolors team blames the static pretrained feature spaces used by existing Frechet losses, which give a fixed and incomplete view of the gap between real and generated distributions.",
					"summary_zh": "AdvFD 针对所谓「Frechet 刷分」问题：生成器的目标 Frechet 指标持续变好，其他特征空间中的视觉质量与分布对齐却停滞甚至恶化。快手可图团队认为症结在于现有 Frechet 损失所依赖的静态预训练特征空间，它对真实与生成分布之间的差异只提供了固定且不完整的视角。",
					"summary_ja": "AdvFD は、目標とするフレシェ距離の指標は改善し続ける一方で、他の特徴空間における視覚品質や分布の一致が停滞・悪化する「フレシェ・ハッキング」に取り組む。Kuaishou の Kolors チームは、既存のフレシェ損失が用いる静的な事前学習済み特徴空間が、実データと生成データの差を固定的かつ不完全にしか捉えられない点に原因を求める。"
				}
			],
			"updated_at": "2026-08-13 13:05 PDT"
		},
		{
			"date": "2026-08-11",
			"items": [
				{
					"rank": 1,
					"title": "SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring",
					"title_zh": "SWE-Bench ProMax：在大规模多语言代码重构上评测智能体",
					"title_ja": "SWE-Bench ProMax: 大規模多言語コードリファクタリングでエージェントを評価する",
					"url": "https://huggingface.co/papers/2608.09802",
					"arxiv_id": "2608.09802",
					"upvotes": 70,
					"first_author": "Yuling Shi",
					"org": "ByteDance",
					"summary": "A ByteDance team proposes SWE-Bench ProMax, a coding-agent benchmark built on large-scale multilingual refactoring instead of bug fixes. The authors cite an audit finding that nearly 60 percent of unsolved SWE-bench Verified instances have flawed tests, and that frontier models can reproduce gold patches verbatim. Refactoring demands behaviour-preserving edits coordinated across many files.",
					"summary_zh": "字节跳动团队提出 SWE-Bench ProMax，这一编码智能体基准以大规模多语言重构任务取代缺陷修复。作者引用一项审计结果：SWE-bench Verified 中未解决实例近 60% 存在测试缺陷，且前沿模型能逐字复现标准补丁。重构要求在众多文件间协同完成保持行为不变的修改。",
					"summary_ja": "ByteDance のチームが、バグ修正ではなく大規模な多言語リファクタリングを土台としたコーディングエージェント向けベンチマーク SWE-Bench ProMax を提案した。著者らは、SWE-bench Verified の未解決事例の約 6 割でテストに欠陥があり、フロンティアモデルが正解パッチをそのまま再現できるとする監査結果を挙げる。リファクタリングは多数のファイルにまたがる挙動保存の編集を要求する。"
				},
				{
					"rank": 2,
					"title": "SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs",
					"title_zh": "SFT 冲突，RL 共存：大模型多任务学习的理论与实证分析",
					"title_ja": "SFT は衝突し、RL は共存する: LLM のマルチタスク学習に関する理論的・実証的分析",
					"url": "https://huggingface.co/papers/2608.03573",
					"arxiv_id": "2608.03573",
					"upvotes": 40,
					"first_author": "Kejian Zhu",
					"org": "Chinese Academic of Science Institute of Automation",
					"summary": "Researchers at the Chinese Academy of Sciences Institute of Automation report that supervised fine-tuning suffers severe task conflicts under multi-stage training while reinforcement learning lets diverse tasks coexist. Traced to the parameter level, RL induces sparse and approximately orthogonal updates across tasks. The paper offers a theoretical account based on multi-task gradient interference.",
					"summary_zh": "中国科学院自动化研究所的研究者发现，在多阶段训练下有监督微调会出现严重的任务冲突，而强化学习则让不同任务得以共存。追溯至参数层面，强化学习在各任务上产生稀疏且近似正交的更新。论文基于多任务梯度干扰给出了理论解释。",
					"summary_ja": "中国科学院自動化研究所の研究者らは、多段階学習では教師ありファインチューニングが深刻なタスク間衝突を起こす一方、強化学習では多様なタスクが共存できると報告した。パラメータ単位で追跡すると、強化学習はタスクごとに疎かつほぼ直交する更新をもたらす。論文はマルチタスクの勾配干渉に基づく理論的説明を与える。"
				},
				{
					"rank": 3,
					"title": "Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning",
					"title_zh": "超越单纯的环境规模扩张：为多模态智能体学习设计有效的环境分布",
					"title_ja": "単なる環境のスケーリングを超えて: マルチモーダルエージェント学習に効く環境分布の設計",
					"url": "https://huggingface.co/papers/2608.03571",
					"arxiv_id": "2608.03571",
					"upvotes": 39,
					"first_author": "Kejian Zhu",
					"org": "Chinese Academic of Science Institute of Automation",
					"summary": "Simply enlarging the pool of multimodal training environments does not always help agents, this group finds. They instead shape the distribution along two axes: Ability-aware Environment Selection picks diverse environment sets, while a difficulty structure controls how hard those environments are. The work argues distribution design matters more than raw environment count.",
					"summary_zh": "该团队发现，单纯扩大多模态训练环境的数量并不总能让智能体受益。他们转而从两个维度调整环境分布：以\"能力感知环境选择\"获取多样的环境集合，并通过难度结构控制环境的难易分布。研究认为分布设计比环境数量本身更为关键。",
					"summary_ja": "マルチモーダルな学習環境をただ増やしてもエージェントが必ず良くなるわけではない、と同グループは指摘する。代わりに分布を二つの軸で設計する。能力を考慮した環境選択で多様な環境集合を得て、難易度の構造でその難しさを制御する。環境の数そのものより分布の設計が重要だと論じている。"
				},
				{
					"rank": 4,
					"title": "Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA",
					"title_zh": "Macaron-V1：迈向自我改进与 Mixture-of-LoRA 的开放持续学习",
					"title_ja": "Macaron-V1: 自己改善と Mixture-of-LoRA による開かれた継続学習に向けて",
					"url": "https://huggingface.co/papers/2608.09819",
					"arxiv_id": "2608.09819",
					"upvotes": 38,
					"first_author": "Mind Lab",
					"org": "Mind Lab",
					"summary": "Mind Lab releases Macaron-V1, an open agent-model family designed to keep learning from real environments after deployment. Adaptation works by recursively improving versioned model-harness pairs, with each configuration evaluated under an external contract to construct its successor. A Mixture-of-LoRA architecture freezes the base model and selects one specialist adapter per user turn.",
					"summary_zh": "Mind Lab 发布 Macaron-V1，这是一个开放的智能体模型家族，目标是在部署之后仍能从真实环境中持续学习。其适应机制通过递归改进带版本的\"模型—框架\"配对实现，每种配置在外部契约下接受评估以构建后继版本。Mixture-of-LoRA 架构冻结基础模型，并在每轮对话中选择一个专家适配器。",
					"summary_ja": "Mind Lab が、配備後も実環境から学び続けることを狙うオープンなエージェントモデル群 Macaron-V1 を公開した。適応はバージョン管理されたモデルとハーネスの組を再帰的に改善する形で進み、各構成は外部の契約のもとで評価され後継の構築に使われる。Mixture-of-LoRA 構成はベースモデルを凍結し、ユーザーのターンごとに専門アダプタを一つ選ぶ。"
				},
				{
					"rank": 5,
					"title": "Motif 3: Technical Report",
					"title_zh": "Motif 3：技术报告",
					"title_ja": "Motif 3: テクニカルレポート",
					"url": "https://huggingface.co/papers/2608.09119",
					"arxiv_id": "2608.09119",
					"upvotes": 23,
					"first_author": "Junghwan Lim",
					"org": "Motif Technologies",
					"summary": "Motif Technologies details Motif 3, a decoder-only Mixture-of-Experts model with 314 billion total parameters and 13.2 billion active per token. Each sparse layer holds 384 routed experts with eight selected per token, giving large expert capacity at limited compute. The architecture centres on Grouped Differential Latent Attention, which pairs grouped differential attention with compressed key-value representations.",
					"summary_zh": "Motif Technologies 详述了 Motif 3：一个仅含解码器的混合专家模型，总参数 3140 亿，每 token 激活 132 亿。每个稀疏层包含 384 个可路由专家，每 token 选取 8 个，在有限算力下提供庞大的专家容量。其架构核心是分组差分潜在注意力，将分组差分注意力与压缩的键值表示结合。",
					"summary_ja": "Motif Technologies が、総パラメータ 3140 億・トークンあたり 132 億を活性化するデコーダ専用 Mixture-of-Experts モデル Motif 3 を解説した。各スパース層は 384 のルーティング専門家を持ち、トークンごとに 8 つを選ぶことで、限られた計算量で大きな専門家容量を確保する。中核となるのは、グループ化差分注意と圧縮された Key-Value 表現を組み合わせた Grouped Differential Latent Attention である。"
				}
			],
			"updated_at": "2026-08-11 01:04 PDT"
		},
		{
			"date": "2026-08-10",
			"items": [
				{
					"rank": 1,
					"title": "StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding",
					"url": "https://huggingface.co/papers/2608.05703",
					"arxiv_id": "2608.05703",
					"upvotes": 14,
					"first_author": "Xichen Zhang",
					"org": "Xiaohongshu",
					"title_zh": "StreamArena：面向连续、交互式与长时程的智能体流式视频理解",
					"title_ja": "StreamArena: 連続的・対話的で長時間にわたるエージェント型ストリーミング動画理解に向けて",
					"summary": "StreamArena is a benchmark for hour-scale interactive video understanding, built from 243 full-length videos averaging 88.8 minutes. The authors argue that current evaluations rely on short clips and multiple-choice formats, which let a baseline reading only the last four frames match far more complex streaming models.",
					"summary_zh": "StreamArena 是一个面向小时级交互式视频理解的基准，由 243 段平均时长 88.8 分钟的完整视频构成。作者指出，现有评测多依赖短片段和多选题形式，以至于仅读取最后四帧的简单基线，就能与复杂得多的流式模型打成平手。",
					"summary_ja": "StreamArena は、平均 88.8 分の全長動画 243 本で構成された、時間単位の対話的動画理解のためのベンチマークである。著者らは、既存の評価が短いクリップと多肢選択形式に依存しているため、最後の 4 フレームしか見ない単純なベースラインでも複雑なストリーミングモデルに匹敵してしまうと指摘する。"
				},
				{
					"rank": 2,
					"title": "DCAS: Decoupling CLI Agent Scaffolding to Internalize Planning across Scaffolds",
					"url": "https://huggingface.co/papers/2608.06113",
					"arxiv_id": "2608.06113",
					"upvotes": 12,
					"first_author": "Kishanthan Thangarajah",
					"org": "Centre for Software Excellence",
					"title_zh": "DCAS：解耦 CLI 智能体脚手架，让规划能力跨脚手架内化",
					"title_ja": "DCAS: CLIエージェントのスキャフォールドを分離し、足場をまたいで計画能力を内在化する",
					"summary": "Open trajectory datasets for command line software engineering agents are collected almost entirely under one scaffold, OpenHands. The authors show that models fine-tuned on that data score well there but degrade under any other scaffold, while untrained base models do not, and trace the gap to learned planning structure.",
					"summary_zh": "用于命令行软件工程智能体的开源轨迹数据集，几乎全部在 OpenHands 这一套脚手架下采集。作者发现，基于这些数据微调的模型在该脚手架上表现良好，换到其他脚手架却明显退化，而未经微调的基座模型不存在这种落差，并将差距归因于被学到的规划结构。",
					"summary_ja": "コマンドライン型ソフトウェア工学エージェント向けの公開軌跡データセットは、そのほとんどが OpenHands という単一のスキャフォールド上で収集されている。著者らは、そのデータで微調整したモデルが同環境では高得点でも他のスキャフォールドでは大きく劣化する一方、未調整のベースモデルにはその差がないことを示し、原因を学習された計画構造に求めている。"
				},
				{
					"rank": 3,
					"title": "When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles",
					"url": "https://huggingface.co/papers/2607.23379",
					"arxiv_id": "2607.23379",
					"upvotes": 10,
					"first_author": "Tobias Bersia",
					"org": "Queen Mary University of London",
					"title_zh": "当激活预言机学会不去读取：微调预言机中与概念相关的盲区",
					"title_ja": "活性化オラクルが「読まない」ことを学ぶとき: ファインチューニングされたオラクルの概念固有の盲点",
					"summary": "Activation oracles are language models trained to answer questions about another model's internal activations. This study argues they are learned systems rather than neutral readouts, and uses a controlled taboo word guessing setup to show concept-specific blind spots shaped by training data and reporting behaviour.",
					"summary_zh": "激活预言机是一类被训练来回答另一模型内部激活状态问题的语言模型。该研究认为它们本身也是被训练出来的系统，而非中立的读数工具，并通过受控的禁忌词猜测实验，展示了由训练数据与汇报行为塑造出的、与特定概念相关的盲区。",
					"summary_ja": "活性化オラクルとは、別のモデルの内部活性化について自然言語で答えるよう訓練された言語モデルである。本研究は、それが中立な読み出しではなく学習されたシステムであると論じ、統制されたタブーワード当ての設定を用いて、訓練データと報告の癖が生む概念固有の盲点を示している。"
				},
				{
					"rank": 4,
					"title": "Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss",
					"url": "https://huggingface.co/papers/2608.03796",
					"arxiv_id": "2608.03796",
					"upvotes": 9,
					"first_author": "Bakbergen Ryskulov",
					"org": "Multiverse Computing",
					"title_zh": "面向大模型的高效知识蒸馏：离线 Top-K logits 与融合分块 KL 损失",
					"title_ja": "LLMのための効率的な知識蒸留: オフラインTop-Kロジットと融合チャンク化KL損失",
					"summary": "Compressed small models are usually recovered through knowledge distillation, an expensive step that largely decides final quality. This practitioner study reports that offline distillation, caching a teacher's top-K logits once and training the student against that cache, matches online distillation at near-identical quality.",
					"summary_zh": "压缩后的小模型通常要靠知识蒸馏来恢复能力，而这一步开销高昂，又在很大程度上决定最终质量。这项面向工程实践的研究表明，把教师模型的 top-K logits 缓存一次再让学生模型据此训练的离线蒸馏，质量几乎与在线蒸馏持平。",
					"summary_ja": "圧縮された小型モデルは通常、知識蒸留によって性能を回復させるが、この工程は高コストでありながら最終的な品質をほぼ左右する。本研究は実務者の視点から、教師モデルの top-K ロジットを一度だけキャッシュして生徒を学習させるオフライン蒸留が、オンライン蒸留とほぼ同等の品質に達すると報告している。"
				},
				{
					"rank": 5,
					"title": "Douyin Multimodal Embedding Model Technical Report",
					"url": "https://huggingface.co/papers/2608.02148",
					"arxiv_id": "2608.02148",
					"upvotes": 7,
					"first_author": "Haonan Chen",
					"org": "ByteDance",
					"title_zh": "抖音多模态嵌入模型技术报告",
					"title_ja": "Douyin マルチモーダル埋め込みモデル技術報告",
					"summary": "ByteDance describes the multimodal embedding model behind search and recommendation on Douyin. The report frames the industrial constraint as needing both efficiency at billion-scale indexing and fine-grained discrimination, which it says contrastive models and chain-of-thought based models each satisfy only in part.",
					"summary_zh": "字节跳动介绍了支撑抖音搜索与推荐的多模态嵌入模型。报告将工业场景的核心约束概括为：既要在十亿级索引规模下保持效率，又要具备细粒度的区分能力，而对比学习模型与基于思维链的模型各自只能满足其中一面。",
					"summary_ja": "バイトダンスが、Douyin の検索・推薦を支えるマルチモーダル埋め込みモデルを解説した技術報告である。産業応用の要件を、十億規模のインデックスでの効率と細粒度の識別能力の両立と整理し、対照学習型モデルと思考連鎖型モデルはそれぞれ一方しか満たせないと論じている。"
				}
			],
			"updated_at": "2026-08-10 13:03 PDT"
		},
		{
			"date": "2026-08-08",
			"items": [
				{
					"rank": 1,
					"title": "AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning",
					"title_zh": "AgentOPSD：面向智能体强化学习的递归自蒸馏",
					"title_ja": "AgentOPSD: エージェント強化学習のための再帰的自己蒸留",
					"url": "https://huggingface.co/papers/2608.05987",
					"arxiv_id": "2608.05987",
					"upvotes": 75,
					"first_author": "Zi-Han Wang",
					"org": "Tsinghua University",
					"summary": "A Tsinghua-led team proposes AgentOPSD, a critic-free method for turn-level credit assignment in agentic reinforcement learning. It aggregates token-level log-probability gaps between a privileged teacher and the student policy, then applies that signal recursively so the few pivotal decisions in long-horizon, multi-turn tasks are credited. It is the day's most upvoted paper with 75 votes.",
					"summary_zh": "清华大学牵头的团队提出 AgentOPSD，一种无需评论家网络、在智能体强化学习中按回合分配信用的方法。该方法先汇总特权教师与学生策略之间的 token 级对数概率差，再以递归方式传递该信号，使长时程多轮任务中真正关键的少数决策获得应有的信用。这是当日票数最高的论文，共 75 票。",
					"summary_ja": "清華大学を中心とするチームは、エージェント強化学習でターン単位のクレジット割り当てを行う critic 不要の手法 AgentOPSD を提案した。特権的な教師と生徒方策のトークン単位の対数確率差を集約し、その信号を再帰的に伝播させることで、長期・多ターンのタスクで結果を左右する少数の決定的な判断を評価できるようにする。本日最多の 75 票を集めた。"
				},
				{
					"rank": 2,
					"title": "WorldClaw: Agentic 3D Open-World Generation at Scale",
					"title_zh": "WorldClaw：大规模智能体式三维开放世界生成",
					"title_ja": "WorldClaw: 大規模なエージェント型3Dオープンワールド生成",
					"url": "https://huggingface.co/papers/2608.05248",
					"arxiv_id": "2608.05248",
					"upvotes": 50,
					"first_author": "Chunchao Guo",
					"org": "Tencent Hunyuan",
					"summary": "Tencent Hunyuan presents WorldClaw, a fully agentic coarse-to-fine framework that builds explorable 3D worlds from open-ended text. Planning agents turn a prompt into a structured specification of regions, terrain, assets, materials and spatial relations, which the system then realizes as a coherent scene with explicit assets suitable for editing and reuse.",
					"summary_zh": "腾讯混元发布 WorldClaw，一个全流程由智能体驱动、从粗到细的框架，可根据开放式文本生成可自由探索的三维世界。规划智能体先把提示词转成关于区域、地形、资产、材质与空间关系的结构化规格，系统再据此构建全局连贯的场景，并输出可编辑、可复用的显式资产。",
					"summary_ja": "テンセント混元は、自由記述のテキストから探索可能な3D世界を構築する、完全にエージェント駆動の粗密フレームワーク WorldClaw を発表した。プランニングエージェントがプロンプトを地域・地形・アセット・マテリアル・空間関係の構造化仕様に変換し、それをもとに全体として整合したシーンを、編集や再利用が可能な明示的アセットとして生成する。"
				},
				{
					"rank": 3,
					"title": "ChronoVision: Temporal Reasoning via Latent State Reconstruction",
					"title_zh": "ChronoVision：通过潜在状态重建实现时序推理",
					"title_ja": "ChronoVision: 潜在状態の再構成による時間的推論",
					"url": "https://huggingface.co/papers/2608.05631",
					"arxiv_id": "2608.05631",
					"upvotes": 33,
					"first_author": "Yifan Shen",
					"org": "PediaMed AI",
					"summary": "ChronoVision targets a known weakness of multimodal LLMs: strong passive perception but poor multi-step temporal reasoning, which the authors trace to the ambiguity of describing continuous visual change in words. During supervised fine-tuning, a reconstructive visual head predicts the latent representation of the final transformed state while an ROI attention module locates the regions that change.",
					"summary_zh": "ChronoVision 针对多模态大模型的一个已知短板：被动感知很强，但多步时序推理很弱，作者将其归因于用语言描述连续视觉变化时的固有歧义。在监督微调阶段，重建式视觉头负责预测最终变换状态的潜在表示，同时 ROI 注意力模块定位发生变化的区域，使视觉逻辑与图像本身对齐。",
					"summary_ja": "ChronoVision は、受動的な知覚には強い一方で多段階の時間的推論が苦手というマルチモーダル LLM の弱点に取り組む。著者らはその原因を、連続的な視覚変化を言語で表すことの曖昧さに求める。教師ありファインチューニングでは、再構成型のビジュアルヘッドが最終的な変換後状態の潜在表現を予測し、ROI アテンションモジュールが変化する領域を特定する。"
				},
				{
					"rank": 4,
					"title": "Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval",
					"title_zh": "从失败中学习：面向统一多模态检索、基于难负例的检索导向思维链",
					"title_ja": "失敗から学ぶ: 統一マルチモーダル検索に向けたハードネガティブによる検索中心の CoT",
					"url": "https://huggingface.co/papers/2608.06060",
					"arxiv_id": "2608.06060",
					"upvotes": 31,
					"first_author": "Zelong Sun",
					"org": "DeepGlint",
					"summary": "DeepGlint researchers build chain-of-thought rationales for multimodal retrieval out of hard negatives rather than the query alone. Existing CoT retrievers explain what the query describes, which leaves vision-language models confusing semantically similar candidates; grounding the reasoning in near-misses surfaces the fine-grained cues that separate a true target from them.",
					"summary_zh": "DeepGlint 的研究者不再只依据查询本身生成思维链，而是用难负例来构造多模态检索的推理过程。现有的思维链检索器只解释查询描述了什么，导致视觉语言模型容易混淆语义相近的候选；把推理建立在这些「差一点就对」的样本上，才能凸显区分真正目标的细粒度线索。",
					"summary_ja": "DeepGlint の研究者らは、クエリだけに基づくのではなく、ハードネガティブからマルチモーダル検索の思考連鎖（CoT）を構築する。従来の CoT 検索器はクエリの内容を説明するにとどまり、視覚言語モデルは意味的に似た候補を取り違えてしまう。惜しくも外れた候補に推論を接地させることで、正解を切り分ける細かな手がかりが浮かび上がる。"
				},
				{
					"rank": 5,
					"title": "From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models",
					"title_zh": "从经济智能体到智能体经济：经济世界模型的系统蓝图",
					"title_ja": "経済エージェントからエージェント経済へ: 経済世界モデルのためのシステム設計図",
					"url": "https://huggingface.co/papers/2608.06020",
					"arxiv_id": "2608.06020",
					"upvotes": 28,
					"first_author": "Jiale Han",
					"org": "FreedomAI",
					"summary": "A team from FreedomAI sets out an implementation roadmap for economic world models, generative systems that simulate an economy from the inside by modeling heterogeneous agents, their beliefs and actions, and the markets and institutions they interact through. The paper organizes such systems into a six-level capability ladder starting from fixed rule-based agents.",
					"summary_zh": "FreedomAI 的团队提出了经济世界模型的实现路线图。这类生成式系统通过刻画异质智能体及其信念与行为，以及他们所处的市场与制度机制，从内部模拟经济的演化。论文把这类系统归纳为一个六级能力阶梯，起点是固定规则驱动的智能体。",
					"summary_ja": "FreedomAI のチームが、経済世界モデルの実装ロードマップを提示した。これは異質なエージェントとその信念・行動、そして相互作用の場となる市場や制度をモデル化し、経済の動きを内側から生成的にシミュレートする仕組みである。論文はこうしたシステムを、固定ルールのエージェントを起点とする6段階の能力階梯として整理している。"
				}
			],
			"updated_at": "2026-08-08 13:04 PDT"
		},
		{
			"date": "2026-08-07",
			"items": [
				{
					"rank": 1,
					"title": "OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models",
					"title_zh": "OSReward：为跨平台计算机操作奖励模型建立标准化评测",
					"title_ja": "OSReward: クロスプラットフォームなコンピュータ操作報酬モデルの標準化評価",
					"url": "https://huggingface.co/papers/2607.28609",
					"arxiv_id": "2607.28609",
					"upvotes": 49,
					"first_author": "Qiushi Sun",
					"org": "NLP Group of The University of Hong Kong",
					"summary": "A 23-author team led from the University of Hong Kong introduces OSReward, a standardized evaluation for the vision-language models used to judge computer-use agent trajectories. The field increasingly relies on VLM judges because neither hand-written verifiers nor human annotators scale, yet whether those judges are reliable has gone largely unexamined.",
					"summary_zh": "由香港大学牵头、共 23 位作者的团队提出 OSReward，为用于判定计算机操作智能体轨迹的视觉语言模型建立标准化评测。由于人工编写的验证器和人工标注都难以规模化，该领域越来越依赖 VLM 充当裁判，但这些裁判是否足够可靠，此前几乎无人系统检验。",
					"summary_ja": "香港大学が主導する23人の著者チームが、コンピュータ操作エージェントの軌跡を判定する視覚言語モデル向けの標準化評価OSRewardを提案した。人手による検証器も人手アノテーションも規模化できないため、この分野はVLMを判定者として使う方向に進んでいるが、その信頼性はこれまでほとんど検証されてこなかった。"
				},
				{
					"rank": 2,
					"title": "Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval",
					"title_zh": "可解释的感知言语 MEG 解码：皮层源与驱动检索的刺激特征",
					"title_ja": "知覚音声のMEG解読を解釈する: 皮質源と検索を駆動する刺激特徴",
					"url": "https://huggingface.co/papers/2608.01481",
					"arxiv_id": "2608.01481",
					"upvotes": 36,
					"first_author": "Ilia Semenkov",
					"summary": "Deep networks can retrieve short segments of heard speech from non-invasive MEG recordings, but their weights do not correspond to any electrophysiological quantity. This paper redesigns both ends of a high-performing retrieval architecture, replacing spatial attention over a flattened sensor layout with spherical harmonics defined on the three-dimensional helmet geometry.",
					"summary_zh": "深度网络可以从无创的脑磁图记录中检索出听到的短段语音，但其权重并不对应任何电生理量。该论文重新设计了一个高性能检索架构的前端与解码器，用定义在三维头盔几何上的球谐函数，取代了在展平传感器布局上做的空间注意力。",
					"summary_ja": "深層ネットワークは非侵襲の脳磁図(MEG)記録から聞き取った短い音声区間を検索できるが、その重みは電気生理学的な量に対応していない。本論文は高性能な検索アーキテクチャのフロントエンドとデコーダの双方を作り替え、平坦化したセンサ配置上の空間アテンションを、三次元ヘルメット形状上で定義した球面調和関数に置き換えた。"
				},
				{
					"rank": 3,
					"title": "GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?",
					"title_zh": "GST-Bench：视觉语言模型能否从视频中建立全局空间意识？",
					"title_ja": "GST-Bench: VLMは動画から全体的な空間認識を獲得できるか",
					"url": "https://huggingface.co/papers/2608.05747",
					"arxiv_id": "2608.05747",
					"upvotes": 30,
					"first_author": "Qifeng Zhang",
					"org": "ByteDance Seed",
					"summary": "ByteDance Seed releases GST-Bench, a video question-answering benchmark for global spatial awareness rather than the local, few-viewpoint perception that existing tests measure. Its human-verified questions are drawn from 6,790 minutes of synthetically generated video and require models to reason about viewpoints never shown in the input.",
					"summary_zh": "字节跳动 Seed 发布 GST-Bench，这是一个针对全局空间意识的视频问答基准，而非现有测试所衡量的单视角或少视角局部感知。其题目经人工核验，取自 6,790 分钟合成生成的视频，要求模型对输入中从未出现过的视角进行推理。",
					"summary_ja": "ByteDance Seedは、既存のテストが測る単一・少数視点の局所的知覚ではなく、全体的な空間認識を問う動画QAベンチマークGST-Benchを公開した。設問は人手で検証され、合成生成された6,790分の動画から作られており、入力に現れない視点についての推論をモデルに求める。"
				},
				{
					"rank": 4,
					"title": "OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents",
					"title_zh": "OneDayAgent：面向自主智能体的长程运行框架",
					"title_ja": "OneDayAgent: 自律エージェントのための長期タスク向けハーネス",
					"url": "https://huggingface.co/papers/2608.05013",
					"arxiv_id": "2608.05013",
					"upvotes": 29,
					"first_author": "Jingsheng Zheng",
					"org": "ZJUNLP",
					"summary": "OneDayAgent is a harness for open-ended everyday requests that run long, cross several environments, and mix modalities. Prior work has attacked goal drift, state loss, and context overflow one at a time; the Zhejiang University group asks whether a single harness can handle all three at once and stay effective across different model backends.",
					"summary_zh": "OneDayAgent 是一个面向开放式日常需求的运行框架，这类任务链条长、跨越多个环境且涉及多种模态。以往工作多是逐一应对目标漂移、状态丢失和上下文溢出，而浙江大学团队要问的是：单一框架能否同时处理这三者，并在不同底层模型上都保持有效。",
					"summary_ja": "OneDayAgentは、長期にわたり複数の環境をまたぎ、複数のモダリティが混在する日常的で自由度の高い依頼に対応するハーネスである。従来研究は目標のずれ、状態の喪失、コンテキスト溢れを個別に扱ってきたが、浙江大学のグループは単一のハーネスで三つを同時に扱い、異なるモデル基盤でも有効性を保てるかを問う。"
				},
				{
					"rank": 5,
					"title": "EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning",
					"title_zh": "EnvACE：通过世界预演内化环境动态的智能体强化学习",
					"title_ja": "EnvACE: 世界リハーサルで環境ダイナミクスを内在化するエージェント強化学習",
					"url": "https://huggingface.co/papers/2608.06197",
					"arxiv_id": "2608.06197",
					"upvotes": 28,
					"first_author": "Zishan Xu",
					"org": "Tencent",
					"summary": "Tencent's EnvACE trains tool-using agents without touching a real or simulated environment. The policy alternates between issuing a tool call and playing the environment that answers it, a loop the authors call world rehearsal, then conditions later decisions on its own rehearsed responses. The target is the cost of building and verifying executable training environments.",
					"summary_zh": "腾讯提出的 EnvACE 在训练调用工具的智能体时，完全不接触真实或模拟环境。策略在发出工具调用与扮演回应该调用的环境之间交替，作者称之为世界预演，随后再基于自己预演出的回应做后续决策。其目标是压低构建与验证可执行训练环境的成本。",
					"summary_ja": "テンセントのEnvACEは、実環境にもシミュレータにも触れずにツール利用エージェントを訓練する。方策はツール呼び出しと、それに応答する環境役を交互に演じ、著者らが世界リハーサルと呼ぶこのループで生成した応答を踏まえて以降の判断を行う。狙いは実行可能な訓練環境の構築と検証にかかるコストの削減である。"
				}
			],
			"updated_at": "2026-08-07 13:06 PDT"
		},
		{
			"date": "2026-08-06",
			"items": [
				{
					"rank": 1,
					"title": "ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment",
					"title_zh": "ABSeeker：通过答案回溯信用分配训练长程搜索智能体",
					"title_ja": "ABSeeker：回答バックトラックによるクレジット割り当てで長期ホライズン検索エージェントを訓練する",
					"url": "https://huggingface.co/papers/2608.05102",
					"arxiv_id": "2608.05102",
					"upvotes": 52,
					"first_author": "Yijun Lu",
					"org": "Shanghai Jiao Tong University",
					"summary": "Researchers at Shanghai Jiao Tong University propose Answer-Backtracked Credit Assignment (ABC), a fine-grained framework for training long-horizon search agents. Rather than treating all steps in a trajectory uniformly during SFT and RL, ABC converts sparse trajectory-level outcomes into per-step credit, separating useful actions from erroneous or redundant ones. It drew 52 upvotes, the day's highest.",
					"summary_zh": "上海交通大学的研究者提出答案回溯信用分配（ABC），一个用于训练长程搜索智能体的细粒度框架。不同于在 SFT 与 RL 中对轨迹内所有步骤一视同仁，ABC 将稀疏的轨迹级结果转化为逐步信用，以区分有用动作与错误或冗余动作。该论文获得 52 个赞，为当日最高。",
					"summary_ja": "上海交通大学の研究者らが、長期ホライズンの検索エージェントを訓練するための細粒度フレームワーク「Answer-Backtracked Credit Assignment（ABC）」を提案。SFTとRLで軌跡内の全ステップを一律に扱うのではなく、疎な軌跡レベルの結果をステップ単位のクレジットに変換し、有用な行動と誤り・冗長な行動を区別する。本論文は当日最多の52アップボートを集めた。"
				},
				{
					"rank": 2,
					"title": "Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes",
					"title_zh": "迈向多模态预训练的物理学：知识流动、模态协同、早期统一与配方",
					"title_ja": "マルチモーダル事前学習の物理へ：知識フロー、モダリティ相乗効果、早期統一、レシピ",
					"url": "https://huggingface.co/papers/2608.05000",
					"arxiv_id": "2608.05000",
					"upvotes": 39,
					"first_author": "Junlin Han",
					"org": "AI at Meta",
					"summary": "Meta researchers run a systematic empirical exploration of natively unified multimodal pretraining. Controlled experiments on synthetic and large-scale real-world datasets yield four insights spanning knowledge flow, modality synergy, early unification, and training recipes, aiming to clarify a design space that unified training has so far left underexplored.",
					"summary_zh": "Meta 的研究者对原生统一的多模态预训练进行了系统的实证探索。在合成与大规模真实数据集上的受控实验得出四条洞见，涵盖知识流动、模态协同、早期统一与训练配方，旨在厘清统一训练中迄今仍未充分探索的设计空间。",
					"summary_ja": "Metaの研究者らが、ネイティブに統一されたマルチモーダル事前学習を体系的に実証研究した。合成データと大規模実データでの統制実験から、知識フロー、モダリティ相乗効果、早期統一、学習レシピにわたる4つの知見を導き、これまで未開拓だった設計空間の解明を目指す。"
				},
				{
					"rank": 3,
					"title": "Quo Vadis, World Modeling?",
					"title_zh": "世界建模，路在何方？",
					"title_ja": "世界モデリングはどこへ向かうのか",
					"url": "https://huggingface.co/papers/2608.02713",
					"arxiv_id": "2608.02713",
					"upvotes": 31,
					"first_author": "Yu Yang",
					"org": "Shanghai AI Laboratory",
					"summary": "A 20-author team centered at Shanghai AI Laboratory takes stock of world modeling for continually improving agents. Arguing that classical future-state prediction is useful but narrow, the paper conceptualizes agent-centric interactive world models that give agents cheaper, more controllable feedback than direct real-environment interaction.",
					"summary_zh": "以上海人工智能实验室为核心的 20 人作者团队梳理了面向持续改进智能体的世界建模。论文认为经典的未来状态预测有用但过于狭窄，进而提出以智能体为中心的交互式世界模型，为智能体提供比直接与真实环境交互更低成本、更可控的反馈。",
					"summary_ja": "上海AIラボを中心とする20名の著者チームが、継続的に改善するエージェントのための世界モデリングを総括。従来の未来状態予測は有用だが狭いとし、実環境との直接対話よりも低コストで制御しやすいフィードバックをエージェントに与える、エージェント中心のインタラクティブ世界モデルを概念化した。"
				},
				{
					"rank": 4,
					"title": "PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents",
					"title_zh": "PAST-Bench：评测个人智能体递归自我改进的基础",
					"title_ja": "PAST-Bench：パーソナルエージェントにおける再帰的自己改善の基盤をベンチマークする",
					"url": "https://huggingface.co/papers/2608.04003",
					"arxiv_id": "2608.04003",
					"upvotes": 28,
					"first_author": "Shuhan Xue",
					"org": "Princeton University",
					"summary": "Princeton researchers introduce PAST-Bench, a benchmark for whether personal AI agents actually improve from retained experience. Agents run ordered sequences of fresh-session tasks under matched conditions that switch retained experience on and off, spanning 26 scenarios, isolating the foundations of recursive self-improvement.",
					"summary_zh": "普林斯顿大学的研究者推出 PAST-Bench，用于检验个人 AI 智能体是否真的能从留存经验中获得提升。智能体在开启与关闭留存经验的匹配条件下，按顺序执行一系列全新会话任务，覆盖 26 个场景，以隔离递归自我改进的基础能力。",
					"summary_ja": "プリンストン大学の研究者らが、パーソナルAIエージェントが蓄積した経験から実際に改善するのかを検証するベンチマーク「PAST-Bench」を発表。エージェントは経験保持のオン・オフを揃えた条件下で新規セッションのタスク列を順に実行し、26のシナリオにわたって再帰的自己改善の基盤を切り分けて測定する。"
				},
				{
					"rank": 5,
					"title": "LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models",
					"title_zh": "LLaDA MoE v2：扩展混合专家扩散语言模型",
					"title_ja": "LLaDA MoE v2：Mixture-of-Experts拡散言語モデルのスケーリング",
					"url": "https://huggingface.co/papers/2608.03457",
					"arxiv_id": "2608.03457",
					"upvotes": 27,
					"first_author": "Fengqi Zhu",
					"org": "GSAI-ML",
					"summary": "LLaDA MoE v2 characterizes how optimization hyperparameters, compute allocation, and architecture scale for mixture-of-experts diffusion language models. The study identifies quantitative differences from autoregressive scaling trends, including an optimal batch size that grows faster and an optimal learning rate that decays more rapidly with compute.",
					"summary_zh": "LLaDA MoE v2 系统刻画了混合专家扩散语言模型在优化超参数、算力分配与架构上的扩展规律。研究发现其与自回归模型已知的扩展趋势存在量化差异，包括最优批大小随算力增长更快，而最优学习率随算力衰减更快。",
					"summary_ja": "LLaDA MoE v2は、Mixture-of-Experts型拡散言語モデルにおける最適化ハイパーパラメータ、計算資源配分、アーキテクチャのスケーリング特性を体系的に明らかにした。自己回帰モデルで報告されてきた傾向との定量的な違いを特定しており、最適バッチサイズは計算量とともにより速く増加し、最適学習率はより急速に減衰するという。"
				}
			],
			"updated_at": "2026-08-06 13:07 PDT"
		},
		{
			"date": "2026-08-05",
			"items": [
				{
					"rank": 1,
					"title": "MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations",
					"title_zh": "MerchantBench：面向电商运营长期一致性的LLM智能体基准",
					"title_ja": "MerchantBench：eコマース運営における長期一貫性を測るLLMエージェントベンチマーク",
					"url": "https://huggingface.co/papers/2607.28956",
					"arxiv_id": "2607.28956",
					"upvotes": 83,
					"first_author": "Qiming Shi",
					"org": "alibaba",
					"summary": "Alibaba researchers present MerchantBench, a benchmark testing whether LLM agents stay coherent over long horizons in seller-side e-commerce operations. Agents run a persistent store where actions constrain future choices and feedback arrives at varying delays, so incoherent decisions accumulate into measurable costs. It targets a gap left by benchmarks built on bounded tasks with immediate success criteria.",
					"summary_zh": "阿里巴巴研究团队提出MerchantBench，用于检验LLM智能体能否在卖家侧电商运营中长期保持连贯而有目的的行为。智能体在持久环境中经营店铺，行动会约束后续选择，反馈到达时间不一，不连贯的决策会累积成可衡量的损失。该基准填补了现有评测以短任务和即时成功标准为主的空白。",
					"summary_ja": "アリババの研究チームは、販売者側のeコマース運営でLLMエージェントが長期にわたり一貫した行動を保てるかを検証するベンチマークMerchantBenchを発表した。エージェントは永続的な環境で店舗を運営し、行動が将来の選択を制約し、フィードバックは様々な遅延で届くため、一貫性を欠く判断は測定可能な損失として蓄積する。短期タスクと即時的な成功基準に偏った既存評価の空白を埋めるものだ。"
				},
				{
					"rank": 2,
					"title": "AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling",
					"title_zh": "AURORA-LM：面向连续潜空间扩散语言建模的自编码统一表征",
					"title_ja": "AURORA-LM：連続潜在空間拡散言語モデリングのための自己符号化統一表現",
					"url": "https://huggingface.co/papers/2608.02602",
					"arxiv_id": "2608.02602",
					"upvotes": 71,
					"first_author": "Jiajun Liang",
					"org": "Nanjing University",
					"summary": "Nanjing University researchers propose AURORA-LM, which generates text with diffusion in a continuous latent space rather than discrete tokens. Instead of compressing latents to ease diffusion, it keeps a high-capacity, decodable text latent and trains the diffusion model to learn its distribution directly, preserving token-level fidelity. The approach brings text closer to how images, video, and audio are modeled.",
					"summary_zh": "南京大学研究者提出AURORA-LM，让文本生成在连续潜空间中通过扩散完成，而非依赖离散token。它不为迁就扩散模型而压缩潜表征，而是保留高容量、可解码的文本潜变量，并让扩散模型直接学习其分布，从而保持token级保真度。该方法使文本建模更接近图像、视频和音频的连续生成范式。",
					"summary_ja": "南京大学の研究者らは、離散トークンではなく連続潜在空間での拡散によりテキストを生成するAURORA-LMを提案した。拡散を容易にするために潜在表現を圧縮するのではなく、高容量で復号可能なテキスト潜在を保持し、その分布を拡散モデルに直接学習させることでトークンレベルの忠実度を保つ。テキスト生成を画像・動画・音声と同様の連続的モデリングに近づける試みである。"
				},
				{
					"rank": 3,
					"title": "Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing",
					"title_zh": "Hunyuan3D-Buffalo 1.0：面向可扩展3D生成、理解与编辑的统一多模态模型",
					"title_ja": "Hunyuan3D-Buffalo 1.0：スケーラブルな3D生成・理解・編集のための統一マルチモーダルモデル",
					"url": "https://huggingface.co/papers/2608.02711",
					"arxiv_id": "2608.02711",
					"upvotes": 70,
					"first_author": "Junliang Ye",
					"org": "Tencent Hunyuan",
					"summary": "Tencent Hunyuan presents Hunyuan3D-Buffalo 1.0, a single architecture unifying 3D understanding, text-to-3D generation, instruction-guided 3D editing, and text-grounded part generation. To overcome the shortage of geometrically consistent editing data, the team built an 87-million-scale 3D multimodal dataset for training. It extends the unified-model trend from images into 3D content creation.",
					"summary_zh": "腾讯混元发布Hunyuan3D-Buffalo 1.0，用单一架构统一3D理解、文本生成3D、指令引导3D编辑和文本定位的部件生成。为解决几何一致的编辑数据稀缺问题，团队构建了8700万规模的3D多模态数据集用于训练。它将图像领域的统一模型趋势延伸到3D内容创作。",
					"summary_ja": "テンセント混元は、3D理解、テキストからの3D生成、指示に基づく3D編集、テキストに紐づくパーツ生成を単一アーキテクチャで統合するHunyuan3D-Buffalo 1.0を発表した。幾何的に一貫した編集データの不足を補うため、8700万規模の3Dマルチモーダルデータセットを構築して学習に用いた。画像で進む統一モデルの潮流を3D制作へと広げる成果だ。"
				},
				{
					"rank": 4,
					"title": "VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation",
					"title_zh": "VAD：多模态在线策略蒸馏中目标重建的视觉证据归因",
					"title_ja": "VAD：マルチモーダル・オンポリシー蒸留における目標再構成の視覚的証拠帰属",
					"url": "https://huggingface.co/papers/2607.28590",
					"arxiv_id": "2607.28590",
					"upvotes": 43,
					"first_author": "Kangning Zhang",
					"summary": "This paper introduces Visual Attribution Distillation for multimodal on-policy distillation, where a privileged teacher corrects student-generated trajectories. Because the teacher's next-token corrections mix visual signals with linguistic priors and teacher-specific effects, VAD uses counterfactual target reconstruction to estimate which part of each correction is actually supported by visual evidence.",
					"summary_zh": "该论文提出视觉归因蒸馏（VAD），用于多模态在线策略蒸馏，即由特权视角教师纠正学生自生成的轨迹。由于教师的逐token纠正混杂了视觉信号、语言先验和教师自身偏差，VAD通过反事实目标重建来估计每次纠正中真正由视觉证据支撑的部分。",
					"summary_ja": "本論文はマルチモーダル・オンポリシー蒸留向けのVisual Attribution Distillation（VAD）を提案する。特権的な視点を持つ教師が生徒の生成軌跡を修正する枠組みだが、その修正には視覚信号に加えて言語的事前知識や教師固有の影響が混在する。VADは反実仮想の目標再構成により、修正のうち視覚的証拠に裏付けられた部分を推定する。"
				},
				{
					"rank": 5,
					"title": "Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent",
					"title_zh": "Video-DeepResearch：迈向下一代多模态深度研究智能体",
					"title_ja": "Video-DeepResearch：次世代マルチモーダル・ディープリサーチエージェントへ",
					"url": "https://huggingface.co/papers/2608.03979",
					"arxiv_id": "2608.03979",
					"upvotes": 43,
					"first_author": "Zhen Fang",
					"summary": "Video-DeepResearch extends multimodal research agents from static images to continuous video streams, pairing dense spatiotemporal grounding with open-web exploration. The authors identify two bottlenecks in current models, a bias toward textual search over visual tools and reliance on parametric memory instead of genuine tool use, and address them with a decoupled perception-exploration pipeline.",
					"summary_zh": "Video-DeepResearch将多模态研究智能体从静态图像扩展到连续视频流，把稠密的时空定位与开放网络探索结合起来。作者指出当前模型的两大瓶颈：偏向文本搜索而绕开视觉工具的模态偏差，以及依赖参数化记忆而非真实工具执行的知识泄漏，并用解耦的感知-探索流水线加以解决。",
					"summary_ja": "Video-DeepResearchはマルチモーダル研究エージェントを静止画像から連続的な動画ストリームへ拡張し、密な時空間グラウンディングとオープンウェブ探索を組み合わせる。著者らは現行モデルの二つのボトルネック、視覚ツールを避けて文字検索に頼るモダリティバイアスと、実際のツール実行ではなく内部記憶に依存する知識漏出を指摘し、知覚と探索を分離したパイプラインで対処する。"
				}
			],
			"updated_at": "2026-08-05 13:07 PDT"
		},
		{
			"date": "2026-08-04",
			"items": [
				{
					"rank": 1,
					"title": "DAPD: Dual-Anchored Policy Distillation",
					"title_zh": "DAPD：双锚定策略蒸馏",
					"title_ja": "DAPD：デュアルアンカー方策蒸留",
					"url": "https://huggingface.co/papers/2608.01735",
					"arxiv_id": "2608.01735",
					"upvotes": 54,
					"first_author": "Jianyu Wu",
					"org": "Shanghai AI Laboratory",
					"summary": "Researchers at Shanghai AI Laboratory identify a privilege illusion in on-policy distillation: students trained against privileged teachers learn behavior they cannot reproduce from their inference-time context. The paper traces the failure to information asymmetry between teacher and student and proposes Dual-Anchored Policy Distillation to resolve it.",
					"summary_zh": "上海人工智能实验室的研究者指出了在线策略蒸馏中的“特权幻觉”：以持有特权信息的教师训练出的学生，会学到无法在推理时上下文中复现的行为。论文将这一失效归因于教师与学生之间的信息不对称，并提出双锚定策略蒸馏（DAPD）来加以解决。",
					"summary_ja": "上海 AI 研究所の研究チームは、オンポリシー蒸留における「特権の幻想」を指摘した。特権情報を持つ教師で訓練された生徒モデルは、推論時のコンテキストからは再現できない挙動を学習してしまう。論文はこの失敗の根本原因を教師と生徒の情報非対称性に求め、解決策としてデュアルアンカー方策蒸留（DAPD）を提案する。"
				},
				{
					"rank": 2,
					"title": "Progressive Agent Skill Generation via Reinforcement Learning",
					"title_zh": "基于强化学习的渐进式智能体技能生成",
					"title_ja": "強化学習による漸進的エージェントスキル生成",
					"url": "https://huggingface.co/papers/2608.01678",
					"arxiv_id": "2608.01678",
					"upvotes": 49,
					"first_author": "Junhao Shen",
					"org": "The Chinese University of Hong Kong - Database Group",
					"summary": "Skill-alpha frames agent skill generation as a reinforcement learning problem instead of relying on hand-designed heuristics or pipeline-style consolidation. Because skills lack a natural supervision signal for relevance or correctness, the method scores them by whether they improve the agent's behavior on downstream tasks.",
					"summary_zh": "Skill-alpha 将智能体技能生成建模为强化学习问题，而非依赖人工设计的启发式规则或流水线式整合。由于技能缺乏基于相关性或正确性的天然监督信号，该方法以技能能否改善智能体在下游任务上的表现来评定其价值。",
					"summary_ja": "Skill-alpha は、人手で設計したヒューリスティクスやパイプライン型の統合に頼らず、エージェントのスキル生成を強化学習の問題として定式化する。スキルには関連性や正しさに基づく自然な教師信号が無いため、下流タスクでエージェントの挙動を改善できるかどうかでスキルの価値を評価する。"
				},
				{
					"rank": 3,
					"title": "Meshy T2: Fast Native Mesh Generation with Flow Matching",
					"title_zh": "Meshy T2：基于流匹配的快速原生网格生成",
					"title_ja": "Meshy T2：フローマッチングによる高速ネイティブメッシュ生成",
					"url": "https://huggingface.co/papers/2607.28675",
					"arxiv_id": "2607.28675",
					"upvotes": 47,
					"first_author": "Jiale Xu",
					"org": "Meshy",
					"summary": "Meshy T2 is a native mesh generation framework built on flow matching, aimed at producing artist-style topology fast enough for interactive 3D asset creation. At its core, a vertex-set mesh VAE encodes a mesh into one continuous latent token per vertex, sidestepping the slow, error-prone autoregressive token decoding of mainstream approaches.",
					"summary_zh": "Meshy T2 是一个基于流匹配的原生网格生成框架，目标是以足够快的速度生成具有美术师风格拓扑的网格，满足交互式 3D 资产创作需求。其核心是一个顶点集网格 VAE，将网格编码为每个顶点一个连续潜在 token，从而绕开主流方法中缓慢且易累积误差的自回归 token 解码。",
					"summary_ja": "Meshy T2 はフローマッチングに基づくネイティブメッシュ生成フレームワークで、インタラクティブな 3D アセット制作に耐える速度でアーティスト品質のトポロジーを生成することを狙う。中核となる頂点集合メッシュ VAE がメッシュを頂点ごとに 1 つの連続潜在トークンへ符号化し、主流手法の遅く誤差が蓄積しやすい自己回帰的トークンデコードを回避する。"
				},
				{
					"rank": 4,
					"title": "UEmbed: Unified Sparse and Dense Multimodal Embeddings",
					"title_zh": "UEmbed：统一的稀疏与稠密多模态嵌入",
					"title_ja": "UEmbed：疎と密を統一したマルチモーダル埋め込み",
					"url": "https://huggingface.co/papers/2608.02583",
					"arxiv_id": "2608.02583",
					"upvotes": 41,
					"first_author": "Tingyu Song",
					"org": "Alibaba-NLP",
					"summary": "UEmbed is a decoder-only multimodal embedding model from Alibaba-NLP that produces both sparse lexical and dense representations in a single causal forward pass. It extends learned sparse retrieval beyond encoder-style bidirectional architectures and drops the auxiliary cross-modal modules earlier multimodal systems leaned on.",
					"summary_zh": "UEmbed 是 Alibaba-NLP 推出的仅解码器多模态嵌入模型，在一次因果前向计算中同时产出稀疏词汇表示和稠密表示。它将学习式稀疏检索从编码器式双向架构中解放出来，并去掉了此前多模态系统所依赖的辅助跨模态模块。",
					"summary_ja": "UEmbed は Alibaba-NLP によるデコーダのみのマルチモーダル埋め込みモデルで、1 回の因果的フォワードパスで疎な語彙表現と密な表現の両方を生成する。学習型スパース検索をエンコーダ型双方向アーキテクチャの制約から解き放ち、従来のマルチモーダルシステムが頼っていた補助的なクロスモーダルモジュールも不要にした。"
				},
				{
					"rank": 5,
					"title": "AISPA: User-Centric System Prompt Auditing for Large Language Model Applications",
					"title_zh": "AISPA：面向大语言模型应用的以用户为中心的系统提示词审计",
					"title_ja": "AISPA：LLM アプリケーションのユーザー中心システムプロンプト監査",
					"url": "https://huggingface.co/papers/2607.28617",
					"arxiv_id": "2607.28617",
					"upvotes": 34,
					"first_author": "Xiangning Lin",
					"org": "Stanford University",
					"summary": "Stanford researchers introduce AISPA, a user-centric framework for systematically auditing the system prompts that govern commercial AI applications. It evaluates parts of a prompt along eight dimensions that matter to users, targeting the trust and accountability gap created by prompts that are rarely disclosed to the public or regulators.",
					"summary_zh": "斯坦福研究者提出 AISPA，一个以用户为中心、系统化审计商业 AI 应用系统提示词的框架。它沿用户关切的八个维度评估提示词的各个部分，旨在弥合系统提示词极少向公众或监管者披露所造成的信任与问责缺口。",
					"summary_ja": "スタンフォード大学の研究チームが AISPA を提案。商用 AI アプリケーションを制御するシステムプロンプトを体系的に監査する、ユーザー中心のフレームワークだ。プロンプトの各部分をユーザーにとって重要な 8 つの観点で評価し、システムプロンプトが一般にも規制当局にもほとんど開示されないことで生じる信頼性と説明責任のギャップに切り込む。"
				}
			],
			"updated_at": "2026-08-04 13:04 PDT"
		},
		{
			"date": "2026-08-03",
			"items": [
				{
					"rank": 1,
					"title": "Mental World Modeling",
					"title_zh": "心智世界建模",
					"title_ja": "メンタル世界モデリング",
					"url": "https://huggingface.co/papers/2607.27201",
					"arxiv_id": "2607.27201",
					"upvotes": 52,
					"first_author": "Hao Fei",
					"org": "Mental World Model",
					"summary": "The paper proposes Mental World Modeling, a framework that makes hidden mental states such as beliefs, desires, intentions, and feelings core variables of a world model. The authors argue that a model tracking only the physical scene predicts the wrong action for the right-looking scene, since human behavior is driven by what agents know and believe rather than by the scene alone.",
					"summary_zh": "论文提出心智世界建模（MWM），一个将信念、欲望、意图和情感等隐藏心智状态作为世界模型核心变量的通用理论框架。作者认为，只跟踪物理场景的模型会在看似正确的场景下预测出错误的行动，因为人类行为由主体的认知与信念驱动，而非仅由场景本身决定。",
					"summary_ja": "本論文は、信念・欲求・意図・感情といった隠れた心的状態を世界モデルの中核変数とする枠組み「メンタル世界モデリング（MWM）」を提案する。物理的シーンだけを追跡するモデルは、見た目には正しいシーンに対して誤った行動を予測してしまうと著者らは論じる。人間の行動はシーンそのものではなく、各主体が何を知り何を信じているかによって駆動されるためである。"
				},
				{
					"rank": 2,
					"title": "N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation",
					"title_zh": "N_0-TWAM：面向富接触操作的触觉原生世界-行动模型规模化",
					"title_ja": "N_0-TWAM：接触の多い操作に向けた触覚ネイティブ世界行動モデルのスケーリング",
					"url": "https://huggingface.co/papers/2607.23783",
					"arxiv_id": "2607.23783",
					"upvotes": 26,
					"first_author": "NeoteAI Team",
					"org": "NeoteAI",
					"summary": "NeoteAI presents N_0-TWAM, a tactile-native world-action model for contact-rich manipulation that predicts both future vision and future contact. It is pre-trained with visuo-tactile joint training on demonstrations spanning six embodiments and 450 tasks, using the force-based NeoForce representation to condition action generation. The authors call it the first such model trained at large scale.",
					"summary_zh": "NeoteAI 发布 N_0-TWAM，一个面向富接触操作的触觉原生世界-行动模型，可同时预测未来视觉和未来接触。模型在覆盖 6 种具身形态、450 项任务的演示数据上进行大规模视觉-触觉联合预训练，并使用基于力的统一触觉表征 NeoForce，将动作生成建立在物理接触信号之上。作者称这是首个大规模训练的触觉世界-行动模型。",
					"summary_ja": "NeoteAI は、接触の多い操作向けに将来の視覚と将来の接触の両方を予測する触覚ネイティブ世界行動モデル N_0-TWAM を発表した。6 種類のエンボディメントと 450 タスクにわたるデモデータで視覚・触覚の共同学習による大規模事前学習を行い、力ベースの統一触覚表現 NeoForce により行動生成を物理的な接触信号に接地させている。著者らはこれを大規模学習された初の触覚世界行動モデルだとしている。"
				},
				{
					"rank": 3,
					"title": "SAF-OPD: Stable Advantage Fusion for On-Policy Distillation",
					"title_zh": "SAF-OPD：面向同策略蒸馏的稳定优势融合",
					"title_ja": "SAF-OPD：オンポリシー蒸留のための安定アドバンテージ融合",
					"url": "https://huggingface.co/papers/2607.29209",
					"arxiv_id": "2607.29209",
					"upvotes": 23,
					"first_author": "Yifan Ding",
					"summary": "The paper combines RL with verifiable rewards, which spreads one sparse reward across all tokens, and on-policy distillation, which scores each token against a stronger teacher but caps performance at teacher quality. The authors show that fusing the two advantages with a fixed coefficient triggers entropy collapse, and propose Stable Advantage Fusion to fix the underlying miscalibrations.",
					"summary_zh": "论文研究如何结合可验证奖励强化学习与同策略蒸馏：前者将单一响应级奖励广播到每个 token，后者用更强教师对每个 token 稠密打分，但性能上限受制于教师水平。作者发现用固定系数融合两种优势会引发熵坍缩，并将其归因于校准失衡，例如 token 级蒸馏优势可能远超有界的 RL 优势并淹没其信号，进而提出稳定的融合方法。",
					"summary_ja": "本論文は、単一の応答レベル報酬を全トークンに一律に与える検証可能報酬付き強化学習（RLVR）と、より強力な教師に対して各トークンを密に採点する一方で性能が教師の水準で頭打ちになるオンポリシー蒸留（OPD）の組み合わせを研究する。固定係数で両者のアドバンテージを融合するとエントロピー崩壊が起きることを発見し、トークンレベルの蒸留アドバンテージが有界な RL アドバンテージを大きく上回り信号をかき消すといった較正のずれに原因を特定した上で、安定した融合手法を提案する。"
				},
				{
					"rank": 4,
					"title": "Scaling Properties of Text Conditioning in Visual Generation",
					"title_zh": "视觉生成中文本条件化的规模化特性",
					"title_ja": "視覚生成におけるテキスト条件付けのスケーリング特性",
					"url": "https://huggingface.co/papers/2607.29679",
					"arxiv_id": "2607.29679",
					"upvotes": 23,
					"first_author": "Zilong Chen",
					"org": "ByteDance Seed",
					"summary": "ByteDance Seed researchers measure how text conditioning scales in visual generation, a question rarely studied because diffusion loss does not scale with prompt token count. They find that converged diffusion loss scales with the amount of structured language in the prompt, decreasing approximately linearly with a likelihood-based measure of prompt structure across controlled training runs.",
					"summary_zh": "字节跳动 Seed 团队测量了视觉生成中文本条件化的经验规模化特性。这一问题此前很少被研究，因为扩散损失并不随提示词 token 数量而变化。他们发现收敛后的扩散损失与提示中结构化语言的数量相关，并用白盒似然指标和黑盒属性指标加以量化：在受控训练中，损失随似然指标近似线性下降。",
					"summary_ja": "ByteDance Seed の研究者らは、視覚生成におけるテキスト条件付けの経験的なスケーリング特性を測定した。拡散損失はプロンプトのトークン数に比例して変化しないため、この問題はこれまでほとんど研究されてこなかった。彼らは、収束後の拡散損失がプロンプト内の構造化言語の量に応じてスケールすることを発見し、ホワイトボックスの尤度指標とブラックボックスの属性指標で定量化した。統制された学習実験では、損失は尤度指標に対しほぼ線形に減少した。"
				},
				{
					"rank": 5,
					"title": "Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants",
					"title_zh": "更少澄清、更好代码：编码助手跨会话个性化歧义适应基准",
					"title_ja": "少ない確認でより良いコードを：コーディングアシスタントのセッション横断的なパーソナライズド曖昧性適応のベンチマーク",
					"url": "https://huggingface.co/papers/2607.26611",
					"arxiv_id": "2607.26611",
					"upvotes": 19,
					"first_author": "Zijian Xu",
					"org": "HKUST",
					"summary": "HKUST researchers formulate personalized ambiguity adaptation as a new task for coding assistants: using a user's resolved session history as memory to disambiguate recurring, user-specific ambiguities in new sessions. Existing methods handle each ambiguous request in isolation, typically by asking clarifying questions, and the benchmark tests whether assistants can instead learn a user's patterns across sessions.",
					"summary_zh": "香港科技大学的研究者将个性化歧义适应定义为编码助手的一项新任务：利用用户已解决的历史会话作为记忆，在新会话中消解反复出现、因人而异的歧义。现有方法通常孤立处理每个含糊请求并依赖追加澄清提问，该基准则检验助手能否跨会话学习用户的个人模式。",
					"summary_ja": "香港科技大学（HKUST）の研究者らは、パーソナライズド曖昧性適応をコーディングアシスタントの新タスクとして定式化した。ユーザーの解決済みセッション履歴をメモリとして活用し、新しいセッションで繰り返し現れるユーザー固有の曖昧さを解消するというものである。既存手法は曖昧なリクエストを個別に扱い、追加の確認質問に頼ることが多いが、このベンチマークはアシスタントがセッションを越えてユーザーのパターンを学習できるかを検証する。"
				}
			],
			"updated_at": "2026-08-03 13:05 PDT"
		},
		{
			"date": "2026-08-01",
			"items": [
				{
					"rank": 1,
					"title": "MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing",
					"title_zh": "MPIE-Bench：解剖学上合理的多人交互编辑基准评测",
					"title_ja": "MPIE-Bench：解剖学的に妥当な複数人物インタラクション編集のベンチマーク",
					"url": "https://huggingface.co/papers/2607.27616",
					"summary": "Researchers introduce MPIE-Bench, a 2,500-sample benchmark for editing multiple named people into shared contact actions such as embracing or carrying. Built from video-mined editing triplets across 405 scenes and 14 interaction categories, it targets failures like fused limbs and interpenetrating bodies that VLM-judge checklists miss but humans spot easily.",
					"summary_zh": "研究者提出 MPIE-Bench，一个包含 2500 个样本的基准，评测把多位指定人物编辑进拥抱、搬抱等身体接触动作的能力。基准由视频挖掘的编辑三元组构成，覆盖 405 个场景和 14 类交互，专门针对肢体融合、身体穿插等 VLM 评审清单会漏掉但人类一眼可见的失败模式。",
					"summary_ja": "研究チームは、複数の特定人物を抱擁や抱き上げなどの接触動作へ編集する能力を評価する 2,500 サンプルのベンチマーク MPIE-Bench を提案した。動画から採掘した編集トリプレットで構成され、405 シーン・14 種類のインタラクションをカバーし、VLM 審査のチェックリストでは見逃されるが人間には明白な、肢体の融合や身体の貫通といった失敗を対象とする。",
					"upvotes": 37,
					"arxiv_id": "2607.27616",
					"first_author": "Jiajia Lin",
					"org": "muset.ai"
				},
				{
					"rank": 2,
					"title": "ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine",
					"title_zh": "ACE-Data-0：以人为中心的环境捕捉作为具身数据引擎",
					"title_ja": "ACE-Data-0：人間中心のアンビエントキャプチャによる身体性データエンジン",
					"url": "https://huggingface.co/papers/2607.28625",
					"summary": "The Ambient Capture Engine (ACE) turns real home environments into spatially calibrated, temporally synchronized recording studios for embodied AI data. It addresses the field's data bottleneck by capturing first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch together as humans pursue goals over time.",
					"summary_zh": "Ambient Capture Engine（ACE）把真实的家居环境改造成空间标定、时间同步的录制工作室，用于采集具身智能数据。它同步记录人类在长时间目标行为中的第一人称感知、全身运动、灵巧操作、物体状态、声音和触觉，以解决该领域的数据瓶颈。",
					"summary_ja": "Ambient Capture Engine（ACE）は、実際の住宅環境を空間的に校正され時間同期された収録スタジオに変え、身体性 AI のデータを収集する。人間が目標を追う過程での一人称視点の知覚、全身運動、巧緻な操作、物体の状態、音、触覚を一体で捉え、この分野のデータボトルネックの解消を狙う。",
					"upvotes": 34,
					"arxiv_id": "2607.28625",
					"first_author": "Yukang Cao"
				},
				{
					"rank": 3,
					"title": "Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation",
					"title_zh": "超越借用的历史：面向交互式角色扮演评测的人格对齐用户模拟",
					"title_ja": "借り物の履歴を超えて：対話型ロールプレイ評価のための人物整合ユーザーシミュレーション",
					"url": "https://huggingface.co/papers/2607.27816",
					"summary": "This paper targets evaluation of role-playing agents, one of the biggest consumer uses of LLMs. Arguing that continuing a fixed dialogue history under a fixed rubric misjudges systems, the authors propose person-aligned user simulation so agents are scored in interactive conversations shaped by the simulated user rather than borrowed transcripts.",
					"summary_zh": "这篇论文关注角色扮演智能体的评测——这是大语言模型最大的消费级应用之一。作者指出，让智能体续写固定对话历史再按固定量规打分会造成误判，因此提出人格对齐的用户模拟，让智能体在由模拟用户塑造的交互式对话中接受评测，而非依赖借来的对话记录。",
					"summary_ja": "本論文は、LLM の最大級のコンシューマー用途であるロールプレイエージェントの評価を扱う。固定された対話履歴の続きを固定ルーブリックで採点する方式では正しく評価できないとし、人物に整合したユーザーシミュレーションを提案。借り物のログではなく、シミュレートされたユーザーが形作る対話の中でエージェントを評価する。",
					"upvotes": 30,
					"arxiv_id": "2607.27816",
					"first_author": "Yuhang Zhu",
					"org": "muset.ai"
				},
				{
					"rank": 4,
					"title": "RefCaptioner: Multi-Reference Image-Grounded Video Captioning",
					"title_zh": "RefCaptioner：多参考图像锚定的视频描述生成",
					"title_ja": "RefCaptioner：複数参照画像に基づく動画キャプション生成",
					"url": "https://huggingface.co/papers/2607.28509",
					"summary": "The Kling team introduces multi-reference image-grounded video captioning, a task requiring factual video descriptions with phrase-level grounding to reference images. RefCaptioner, a two-stage post-training framework combining mixed-data SFT with Hierarchical Coverage-Discounted GRPO, improves reference selection, phrase binding, and distractor rejection.",
					"summary_zh": "可灵团队提出多参考图像锚定的视频描述任务，要求生成的视频描述在短语级别与参考图像对应且符合事实。RefCaptioner 是一个两阶段后训练框架，将混合数据 SFT 与分层覆盖折扣 GRPO 相结合，提升了参考选择、短语绑定和干扰项排除能力。",
					"summary_ja": "Kling チームは、参照画像へのフレーズ単位のグラウンディングを伴う事実に即した動画記述を求める新タスク「複数参照画像に基づく動画キャプション生成」を提案した。RefCaptioner は混合データ SFT と階層的カバレッジ割引 GRPO を組み合わせた2段階のポストトレーニング枠組みで、参照の選択、フレーズの対応付け、ディストラクタの排除を改善する。",
					"upvotes": 25,
					"arxiv_id": "2607.28509",
					"first_author": "Tengfei Liu",
					"org": "Kling Team"
				},
				{
					"rank": 5,
					"title": "See2Think: Do Multimodal Models Really Use Intermediate Visual States?",
					"title_zh": "See2Think：多模态模型真的在利用中间视觉状态吗？",
					"title_ja": "See2Think：マルチモーダルモデルは中間視覚状態を本当に活用しているのか",
					"url": "https://huggingface.co/papers/2607.26769",
					"summary": "See2Think asks whether multimodal models that sketch, annotate, and generate intermediate images during reasoning actually rely on those visual states. The framework pairs See2ThinkBench, 1,200 open-ended visual tasks, with a Visual Action-of-Thought protocol that diagnoses how intermediate visual states are generated, rendered, and used rather than only scoring final answers.",
					"summary_zh": "See2Think 追问一个问题：在推理中绘制草图、做标注、生成中间图像的多模态模型，是否真的依赖这些视觉状态。该框架由包含 1200 个开放式视觉任务的 See2ThinkBench 和视觉思维行动（VAoT）协议组成，不只评判最终答案，而是诊断中间视觉状态如何被生成、渲染和使用。",
					"summary_ja": "See2Think は、推論中にスケッチや注釈、中間画像を生成するマルチモーダルモデルが、その視覚状態に本当に依存しているのかを問う。1,200 件のオープンエンドな視覚タスクからなる See2ThinkBench と、最終回答の採点にとどまらず中間視覚状態の生成・描画・利用を診断する Visual Action-of-Thought プロトコルを組み合わせた評価フレームワークだ。",
					"upvotes": 22,
					"arxiv_id": "2607.26769",
					"first_author": "Siyu Yan"
				}
			],
			"updated_at": "2026-08-01 13:03 PDT"
		},
		{
			"date": "2026-07-31",
			"items": [
				{
					"rank": 1,
					"title": "Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents",
					"title_zh": "Qwen-UI-Agent 技术报告：迈向下一代面向真实世界的基础 GUI 智能体",
					"title_ja": "Qwen-UI-Agent 技術報告：次世代の実世界中心の基盤 GUI エージェントに向けて",
					"url": "https://huggingface.co/papers/2607.28227",
					"summary": "TongyiLab presents Qwen-UI-Agent, a real-world centric foundation GUI agent spanning mobile, computer-use, web, and DeepSearch environments. The report targets agents that operate reliably on real devices, combine GUI interaction with CLI execution, and complete long-horizon tasks while improving with minimal human effort. It leads the day's Hugging Face papers with 264 upvotes.",
					"summary_zh": "通义实验室发布 Qwen-UI-Agent，一个面向真实世界的基础 GUI 智能体，覆盖移动端、计算机操作、网页和 DeepSearch 环境。报告的目标是让智能体在真实设备上可靠运行、将 GUI 交互与 CLI 执行结合、完成长程任务，并以最少的人工投入自我提升。该论文以 264 个赞领跑当日 Hugging Face 论文榜。",
					"summary_ja": "TongyiLab は、モバイル、コンピュータ操作、ウェブ、DeepSearch の各環境にまたがる実世界中心の基盤 GUI エージェント Qwen-UI-Agent を発表した。実機で確実に動作し、GUI 操作と CLI 実行を組み合わせ、長期タスクを完遂しながら最小限の人手で自己改善するエージェントを目指す。264 の賛成票で当日の Hugging Face 論文で首位に立った。",
					"upvotes": 264,
					"arxiv_id": "2607.28227",
					"first_author": "Hanzhang Zhou",
					"org": "TongyiLab"
				},
				{
					"rank": 2,
					"title": "Metis: Memory Foundation Model",
					"title_zh": "Metis：记忆基础模型",
					"title_ja": "Metis：メモリ基盤モデル",
					"url": "https://huggingface.co/papers/2607.26760",
					"summary": "MemTensor introduces memory foundation models, which build native memory capability directly into foundation models instead of relying on external memory modules. The work formalizes native memory in part as a persistent, dynamically evolving memory state, taking a first step toward agents that remember by design rather than by bolted-on retrieval.",
					"summary_zh": "MemTensor 提出记忆基础模型，将原生记忆能力直接构建进基础模型，而不是依赖外部记忆模块。这项工作将原生记忆部分形式化为一种持久且动态演化的记忆状态，朝着让智能体凭设计本身而非外挂检索来记忆迈出第一步。",
					"summary_ja": "MemTensor は、外部メモリモジュールに頼らず基盤モデル自体にネイティブな記憶能力を組み込む「メモリ基盤モデル」を提唱した。ネイティブメモリを持続的かつ動的に進化する記憶状態などとして形式化し、後付けの検索ではなく設計そのものによって記憶するエージェントへの第一歩を示している。",
					"upvotes": 213,
					"arxiv_id": "2607.26760",
					"first_author": "Zeyu Zhang",
					"org": "MemTensor"
				},
				{
					"rank": 3,
					"title": "AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis",
					"title_zh": "AskChem：面向化学文献综合的以论断为中心的基础设施",
					"title_ja": "AskChem：化学文献統合のためのクレーム中心インフラ",
					"url": "https://huggingface.co/papers/2607.28618",
					"summary": "NYU researchers present AskChem, a claim-centered infrastructure for cross-paper chemistry search. Instead of returning ranked document lists, it converts each paper into atomic, typed claims grounded by a source DOI, so scientists and AI agents can assemble cross-paper answers with verifiable provenance.",
					"summary_zh": "纽约大学的研究者提出 AskChem，一个以论断为中心的跨论文化学检索基础设施。它不再返回按相关性排序的文献列表，而是把每篇论文转化为带来源 DOI 的原子化、类型化论断，让科学家和 AI 智能体能够组装出来源可验证的跨论文答案。",
					"summary_ja": "ニューヨーク大学の研究者らは、論文横断の化学検索のためのクレーム中心インフラ AskChem を発表した。ランク付けした文献リストを返す代わりに、各論文を出典 DOI に紐づいた原子的で型付きのクレームへ変換し、科学者や AI エージェントが出所を検証できる論文横断の回答を組み立てられるようにする。",
					"upvotes": 199,
					"arxiv_id": "2607.28618",
					"first_author": "Bing Yan",
					"org": "New York University"
				},
				{
					"rank": 4,
					"title": "PhiZero: A World Model Built Around Physical Language",
					"title_zh": "PhiZero：围绕物理语言构建的世界模型",
					"title_ja": "PhiZero：物理言語を中心に構築された世界モデル",
					"url": "https://huggingface.co/papers/2607.28624",
					"summary": "PhiZero is a physical world model built around physical language, a compact discrete representation of world-state transitions learned from in-the-wild videos via self-supervision. Rather than predicting future video directly in pixel space, it uses this representation to reason explicitly about how the physical world evolves.",
					"summary_zh": "PhiZero 是一个围绕物理语言构建的物理世界模型——物理语言是对世界状态转移的紧凑离散表示，通过自监督从真实场景视频中学得。它不在像素空间直接预测未来视频，而是利用这一表示对物理世界的演化进行显式推理。",
					"summary_ja": "PhiZero は「物理言語」を中心に構築された物理世界モデルである。物理言語とは世界状態の遷移を表す簡潔な離散表現で、実世界の動画から自己教師あり学習で獲得される。ピクセル空間で未来の映像を直接予測するのではなく、この表現を用いて物理世界の変化を明示的に推論する。",
					"upvotes": 148,
					"arxiv_id": "2607.28624",
					"first_author": "Shuyao Shang"
				},
				{
					"rank": 5,
					"title": "Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering",
					"title_zh": "Frontis-MA1：训练面向机器学习工程递归自我改进的 AI4AI 模型",
					"title_ja": "Frontis-MA1：機械学習エンジニアリングにおける再帰的自己改善を目指す AI4AI モデルの訓練",
					"url": "https://huggingface.co/papers/2607.28568",
					"summary": "Frontis AI post-trains Frontis-MA1, a 35B meta-evolution agent for machine learning engineering, on OpenMLE, an open full-stack system for recursive self-improvement research. The stack spans verifiable task environments with execution feedback, operator learning, and long-horizon search, treating MLE as a concrete testbed for AI that improves the process of building AI.",
					"summary_zh": "Frontis AI 在 OpenMLE 上后训练了 Frontis-MA1，一个面向机器学习工程的 35B 元进化智能体；OpenMLE 是用于递归自我改进研究的开放全栈系统。该体系涵盖带执行反馈的可验证任务环境、算子学习和长程搜索，把机器学习工程当作 AI 改进 AI 构建过程的具体试验场。",
					"summary_ja": "Frontis AI は、再帰的自己改善研究のためのオープンなフルスタックシステム OpenMLE 上で、機械学習エンジニアリング向けの 35B メタ進化エージェント Frontis-MA1 をポストトレーニングした。実行フィードバック付きの検証可能なタスク環境、オペレータ学習、長期探索を備え、機械学習エンジニアリングを AI が AI 構築のプロセスを改善するための具体的なテストベッドと位置づけている。",
					"upvotes": 121,
					"arxiv_id": "2607.28568",
					"first_author": "Junlin Yang",
					"org": "Frontis AI"
				}
			],
			"updated_at": "2026-07-31 13:05 PDT"
		},
		{
			"date": "2026-07-30",
			"items": [
				{
					"rank": 1,
					"title": "TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM",
					"title_zh": "TurboVLA：在 RTX 4090 上以 32 Hz、不足 1 GB 显存实时运行的视觉-语言-动作模型",
					"title_ja": "TurboVLA - RTX 4090 上で 32 Hz、1 GB 未満の VRAM で動く実時間 Vision-Language-Action モデル",
					"url": "https://huggingface.co/papers/2607.27205",
					"upvotes": 118,
					"arxiv_id": "2607.27205",
					"first_author": "Hengyi Xie",
					"org": "H-EmbodVis",
					"summary": "Vision-language-action models usually route what the robot sees through a language model before decoding an action, which costs compute and memory on every policy call. TurboVLA reworks that pathway, and the authors report control running at 32 Hz on a single RTX 4090 inside 1 GB of VRAM.",
					"summary_zh": "视觉-语言-动作模型通常先把机器人看到的画面投射进语言模型的表示空间，再解码成动作，这让每一次策略调用都付出可观的算力与显存代价。TurboVLA 重构了这条路径，作者报告在单张 RTX 4090 上以 32 Hz、1 GB 显存以内完成控制。",
					"summary_ja": "Vision-Language-Action モデルは通常、ロボットが見た映像をいったん言語モデルの表現空間に写してから行動へとデコードするため、方策を呼び出すたびに計算量とメモリを大きく消費する。TurboVLA はこの経路を組み替え、単一の RTX 4090 上で 1 GB 未満の VRAM、32 Hz での制御を報告している。"
				},
				{
					"rank": 2,
					"title": "CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization",
					"title_zh": "CoRT：面向 token 级评分标准引导策略优化的反事实回放",
					"title_ja": "CoRT - トークン単位のルーブリック誘導型方策最適化のための反実仮想リプレイ",
					"url": "https://huggingface.co/papers/2607.25659",
					"upvotes": 73,
					"arxiv_id": "2607.25659",
					"first_author": "Bo-Wen Zhang",
					"org": "ByteDance",
					"summary": "Rubric-based reinforcement learning grades model outputs against explicit written criteria, but GRPO-style pipelines flatten those judgments into a single response-level number. CoRT uses counterfactual replay to push the rubric signal down to individual tokens.",
					"summary_zh": "基于评分标准的强化学习，会按明确写出的准则来评判模型输出；但在 GRPO 式流水线中，这些结构化判断最终被压成一个响应级的标量奖励。CoRT 用反事实回放把评分信号下沉到单个 token 层面。",
					"summary_ja": "ルーブリックに基づく強化学習は、明示的に書かれた基準に照らしてモデル出力を評価する。しかし GRPO 型のパイプラインでは、その構造化された判断が応答単位のスカラー報酬へと押し潰されてしまう。CoRT は反実仮想リプレイによって、ルーブリックの信号を個々のトークンにまで届かせる。"
				},
				{
					"rank": 3,
					"title": "HumanCLAW: Can Vision-Language Models Act Through a Body?",
					"title_zh": "HumanCLAW：视觉-语言模型能否通过一具身体去行动？",
					"title_ja": "HumanCLAW - 視覚言語モデルは身体を通じて行動できるか",
					"url": "https://huggingface.co/papers/2607.27180",
					"upvotes": 65,
					"arxiv_id": "2607.27180",
					"first_author": "Siyao Li",
					"org": "Meta Research",
					"summary": "Judging whether a vision-language model can act through a physical body is hard because the outcome mixes the model's decision with motor control, so a failed task does not say which part went wrong. HumanCLAW is built to separate the two.",
					"summary_zh": "要判断一个视觉-语言模型能否通过物理身体行动很困难：结果同时耦合了模型的决策与运动控制，任务失败时很难分清是哪一环出了问题。HumanCLAW 正是为把这两者拆开而设计的。",
					"summary_ja": "視覚言語モデルが身体を通じて行動できるかを評価するのは難しい。結果にはモデルの判断と運動制御の両方が絡むため、課題に失敗してもどちらが原因かを切り分けられないからだ。HumanCLAW はこの二つを分離するために設計されている。"
				},
				{
					"rank": 4,
					"title": "DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator",
					"title_zh": "DecoEvo：求解器与评分标准生成器的分数解耦协同演化",
					"title_ja": "DecoEvo - ソルバーとルーブリック生成器のスコア分離型共進化",
					"url": "https://huggingface.co/papers/2607.25675",
					"upvotes": 54,
					"arxiv_id": "2607.25675",
					"first_author": "Jiangwang Chen",
					"org": "Qwen Business Unit",
					"summary": "Text-space optimization adapts a language model by editing external natural-language artifacts rather than weights, which keeps what changed readable and lets the model stay a black box. DecoEvo evolves the solver and the rubric generator together, with their scores decoupled.",
					"summary_zh": "文本空间优化不改权重，而是通过编辑外部的自然语言产物来适配大模型 —— 好处是改了什么始终可读，模型本身可以当作黑箱。DecoEvo 让求解器与评分标准生成器协同演化，并把两者的分数解耦。",
					"summary_ja": "テキスト空間での最適化は、重みではなく外部の自然言語による生成物を編集してモデルを適応させる。何を変えたかが読める形で残り、モデル自体はブラックボックスのまま扱える点が利点だ。DecoEvo はソルバーとルーブリック生成器を共進化させ、両者のスコアを切り離す。"
				},
				{
					"rank": 5,
					"title": "CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge",
					"title_zh": "CLBench-V：从视觉定位到知识层面评估多模态上下文学习",
					"title_ja": "CLBench-V - グラウンディングから知識までを対象としたマルチモーダル文脈学習の評価",
					"url": "https://huggingface.co/papers/2607.25294",
					"upvotes": 40,
					"arxiv_id": "2607.25294",
					"first_author": "Lai Wei",
					"summary": "Real tasks often require a model to learn from the context it is given rather than lean only on what it absorbed in pre-training. CLBench-V evaluates that ability across multimodal settings, spanning visual grounding through to knowledge-level questions.",
					"summary_zh": "现实任务往往要求模型从当下给定的上下文中学习，而不是只依赖预训练时吸收的知识。CLBench-V 在多模态场景下评估这项能力，覆盖从视觉定位到知识层面的问题。",
					"summary_ja": "現実の課題では、事前学習で得た知識だけに頼るのではなく、与えられた文脈から学ぶ能力がしばしば求められる。CLBench-V はこの能力をマルチモーダルな設定で評価し、視覚的なグラウンディングから知識レベルの問いまでを対象とする。"
				}
			],
			"updated_at": "2026-07-30 12:36 PDT"
		},
		{
			"date": "2026-07-29",
			"items": [
				{
					"rank": 1,
					"title": "HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone",
					"title_zh": "HiFi-UMI：仅用高保真 UMI 数据学习可部署的操作策略",
					"title_ja": "HiFi-UMI: 高忠実度の UMI データだけで実運用可能な操作方策を学習する",
					"url": "https://huggingface.co/papers/2607.25895",
					"arxiv_id": "2607.25895",
					"upvotes": 126,
					"first_author": "Simple AI",
					"org": "Simple World Lab",
					"summary": "Deployable manipulation policies are held back by data that is either accurate or scalable, rarely both: teleoperation is precise but costly, while robot-free UMI capture scales and is usually relegated to pre-training with a small real-robot anchor added later. The authors ask whether raising UMI fidelity can remove that anchor entirely.",
					"summary_zh": "可部署的操作策略受限于数据「要么精确、要么可扩展」的困境：真机遥操作精确但扩展成本高，无机器人的 UMI 采集容易扩展，却通常只用于预训练，再补一小段真机数据做锚定。作者提出的问题是：把 UMI 数据的保真度提上去，能否彻底去掉这个锚。",
					"summary_ja": "実運用可能な操作方策は、正確さと拡張性を兼ね備えたデータの不足に阻まれてきた。実機の遠隔操作は正確だが規模を出しにくく、ロボット不要の UMI 収集は拡張しやすい一方で主に事前学習に回され、後から少量の実機データで補正される。著者らは UMI データの忠実度を上げればその補正を不要にできるかを問う。"
				},
				{
					"rank": 2,
					"title": "A New Role for Relevance: Guiding Corpus Interaction in Agentic Search",
					"title_zh": "相关性的新角色：引导智能体搜索中的语料交互",
					"title_ja": "関連性の新たな役割: エージェント検索におけるコーパス操作を導く",
					"url": "https://huggingface.co/papers/2607.24223",
					"arxiv_id": "2607.24223",
					"upvotes": 79,
					"first_author": "Jiangnan Li",
					"org": "Tencent",
					"summary": "Retrieval agents use relevance to pick top-k content, but the authors argue relevance alone cannot localize, compose, or verify the evidence a complex question needs. Direct Corpus Interaction supports those finer operations through grep-style exploration, yet its relevance-agnostic search surfaces useful clues late and delays convergence.",
					"summary_zh": "检索智能体用相关性挑出 top-k 内容，但作者指出，仅靠相关性无法定位、组合与验证复杂问题所需的证据。Direct Corpus Interaction 用类似 grep 的探索支持这些更细的操作，可它的搜索不看相关性，导致有用线索出现得太晚、收敛被拖慢。",
					"summary_ja": "検索エージェントは関連性で上位 k 件を選ぶが、著者らは関連性だけでは複雑な問いに必要な証拠を特定・構成・検証できないと論じる。Direct Corpus Interaction は grep 的な探索でそうした細かい操作を可能にするが、関連性を考慮しない探索のため有用な手がかりが遅れて現れ、収束が滞る。"
				},
				{
					"rank": 3,
					"title": "StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents",
					"title_zh": "StateAct：面向长周期计算机操作智能体，先看程序状态而非像素",
					"title_ja": "StateAct: 長期的なコンピュータ操作エージェントのために、ピクセルより先にプログラム状態を",
					"url": "https://huggingface.co/papers/2607.22798",
					"arxiv_id": "2607.22798",
					"upvotes": 55,
					"first_author": "Yan Yang",
					"org": "Salesforce AI Research",
					"summary": "Computer-use agents are usually improved by sharpening perception of screenshots. The authors point out that a screenshot is a lossy rendering of the underlying program state - files, application backends, the DOM - and that different states can render to identical pixels, so they condition the agent on that state instead.",
					"summary_zh": "计算机操作智能体通常靠强化对截图的感知来改进。作者指出，截图只是底层程序状态（文件、应用后端、DOM）的有损呈现，不同状态可能渲染出完全相同的像素，因此他们改为让智能体直接基于程序状态决策。",
					"summary_ja": "コンピュータ操作エージェントは通常、スクリーンショットの知覚を強化して改善される。著者らは、スクリーンショットは背後のプログラム状態（ファイル、アプリのバックエンド、DOM）の劣化した描画にすぎず、異なる状態が同一のピクセルになり得ると指摘し、代わりにその状態を条件として与える。"
				},
				{
					"rank": 4,
					"title": "ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition",
					"title_zh": "ReDesign：通过智能体分解从图像中还原可编辑的设计结构",
					"title_ja": "ReDesign: エージェント的分解によって画像から編集可能なデザイン構造を復元する",
					"url": "https://huggingface.co/papers/2607.25565",
					"arxiv_id": "2607.25565",
					"upvotes": 54,
					"first_author": "Jooyeol Yun",
					"org": "KAIST AI",
					"summary": "Turning a raster image back into an editable design file is a costly bottleneck, because editability depends on recovering typography, vector geometry, colors, grouping, and layer ordering together. ReDesign is an agentic framework that grows an editable layer hierarchy by selecting and composing specialized tools across those modalities.",
					"summary_zh": "把位图还原成可编辑的设计文件是设计流程中代价高昂的瓶颈，因为可编辑性取决于同时还原字体排印、矢量几何、颜色、分组与图层顺序。ReDesign 是一个智能体框架，通过选择并组合跨模态的专用工具，逐步生长出可编辑的图层层级。",
					"summary_ja": "ラスタ画像を編集可能なデザインファイルに戻す作業は費用のかかるボトルネックである。編集可能性はタイポグラフィ、ベクタ形状、色、グループ化、レイヤー順序をまとめて復元できるかに依存するからだ。ReDesign は各モダリティの専用ツールを選択・合成しながら編集可能なレイヤー階層を育てるエージェント的枠組みである。"
				},
				{
					"rank": 5,
					"title": "CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents",
					"title_zh": "CodeNib：向编程智能体提供仓库上下文的多视图数据系统",
					"title_ja": "CodeNib: コーディングエージェントにリポジトリ文脈を供給するマルチビュー・データシステム",
					"url": "https://huggingface.co/papers/2607.25431",
					"arxiv_id": "2607.25431",
					"upvotes": 34,
					"first_author": "Zhongming Yu",
					"org": "SysEvol AI Research",
					"summary": "Coding agents keep rediscovering the same repository context because indexes, language servers, and task-local histories are disconnected. CodeNib builds lexical, dense, and structural views per commit, maps results to repository-relative source ranges, and serves ranked search, symbol navigation, and bounded context from one runtime.",
					"summary_zh": "编程智能体反复重新发现同一份仓库上下文，原因是索引、语言服务器与任务内历史彼此割裂。CodeNib 按每次提交构建词法、稠密与结构三种视图，把结果映射回仓库相对的源码区间，并从同一个运行时提供排序检索、符号跳转与受限上下文。",
					"summary_ja": "コーディングエージェントは、インデックス・言語サーバー・タスク内の履歴が分断されているために、同じリポジトリ文脈を何度も探し直す。CodeNib はコミットごとに字句・密ベクトル・構造の各ビューを構築し、結果をリポジトリ相対のソース範囲に対応づけ、ランク付き検索・シンボル移動・上限付き文脈を単一のランタイムから提供する。"
				}
			],
			"updated_at": "2026-07-29 12:40 PDT"
		},
		{
			"date": "2026-07-28",
			"items": [
				{
					"rank": 1,
					"title": "Kimi K3: Open Frontier Intelligence",
					"title_zh": "Kimi K3：开放的前沿智能",
					"title_ja": "Kimi K3: オープンなフロンティア知能",
					"url": "https://huggingface.co/papers/2607.24653",
					"arxiv_id": "2607.24653",
					"upvotes": 250,
					"first_author": "Kimi Team",
					"org": "Moonshot AI",
					"summary": "The report introduces Kimi K3, a 2.8T parameter mixture-of-experts model with 104 billion activated parameters, native vision, and a 1-million-token context window. Alongside Kimi Delta Attention and Attention Residuals, a Stable LatentMoE design activates 16 of 896 routed experts per token, which the team credits for roughly 2.5x better scaling efficiency.",
					"summary_zh": "报告介绍 Kimi K3：2.8 万亿参数的混合专家模型，激活参数 1040 亿，具备原生视觉能力与 100 万 token 上下文窗口。除 Kimi Delta Attention 与 Attention Residuals 外，Stable LatentMoE 让每个 token 只激活 896 个路由专家中的 16 个，团队将约 2.5 倍的扩展效率提升归功于这些设计。",
					"summary_ja": "Kimi K3 を紹介する報告。2.8 兆パラメータの Mixture-of-Experts で、活性化パラメータは 1040 億、ネイティブな視覚能力と 100 万トークンの文脈長を持つ。Kimi Delta Attention と Attention Residuals に加え、Stable LatentMoE が 896 のルーテッドエキスパートのうち 16 をトークンごとに活性化し、これらが約 2.5 倍のスケーリング効率向上をもたらしたとする。"
				},
				{
					"rank": 2,
					"title": "JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents",
					"title_zh": "JarvisHub：面向画布原生多模态创作智能体的开放框架",
					"title_ja": "JarvisHub: キャンバスネイティブなマルチモーダル創作エージェントのためのオープンハーネス",
					"url": "https://huggingface.co/papers/2607.23588",
					"arxiv_id": "2607.23588",
					"upvotes": 101,
					"first_author": "Yunlong Lin",
					"summary": "Creative AI is moving from single-step asset generation toward long-horizon production, the authors argue. Real creative work involves references, drafts, alternatives, edits, failed attempts, version relations, tool actions, and human feedback, which together form an evolving project state that isolated prompt-output exchanges cannot hold.",
					"summary_zh": "作者认为，创意 AI 正从单步素材生成走向长周期生产。真实创作包含参考、草稿、备选方案、修改、失败尝试、版本关系、工具操作与人类反馈，它们共同构成一份不断演进的项目状态，而孤立的提示与输出承载不了这些。",
					"summary_ja": "著者らは、クリエイティブ AI が単発のアセット生成から長期的な制作へ移りつつあると論じる。実際の制作には参照、下書き、代替案、修正、失敗した試み、バージョンの関係、ツール操作、人間のフィードバックが含まれ、それらが進化するプロジェクト状態を形づくる。孤立したプロンプトと出力の往復では、その状態は保持できない。"
				},
				{
					"rank": 3,
					"title": "From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search",
					"title_zh": "从闭源到开源：用多智能体协议蒸馏弥合智能体搜索中的分布差距",
					"title_ja": "プロプライエタリからオープンソースへ: エージェント検索におけるマルチエージェント・プロトコル蒸留で分布ギャップを埋める",
					"url": "https://huggingface.co/papers/2607.24280",
					"arxiv_id": "2607.24280",
					"upvotes": 64,
					"first_author": "Junlin Liu",
					"summary": "Agentic search interleaves multi-step reasoning with retrieval, but outcome-based reinforcement learning gives it only sparse supervision. Proprietary models make strong teachers for denser guidance, yet their hidden logits rule out conventional logit-matching, so the authors distill at the level of a multi-agent protocol instead.",
					"summary_zh": "智能体搜索把多步推理与检索交错进行，但以结果为准的强化学习只能提供稀疏监督。闭源强模型是理想的教师，可它们不公开 logits，常规的 logit 匹配无从谈起；因此作者改在多智能体协议的层面做蒸馏。",
					"summary_ja": "エージェント検索は多段推論と検索を交互に行うが、結果ベースの強化学習では監督が疎になる。プロプライエタリモデルは有力な教師だが logits が非公開のため従来の logit マッチングは使えない。そこで著者らはマルチエージェント・プロトコルの水準で蒸留する。"
				},
				{
					"rank": 4,
					"title": "Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation",
					"title_zh": "重新审视 on-policy 扩散蒸馏中的无分类器引导",
					"title_ja": "オンポリシー拡散蒸留における classifier-free guidance の再考",
					"url": "https://huggingface.co/papers/2607.24731",
					"arxiv_id": "2607.24731",
					"upvotes": 61,
					"first_author": "Bingnan Li",
					"summary": "On-policy distillation adapts a diffusion model by querying a teacher along the student's own trajectories, but its behaviour under classifier-free guidance is poorly understood. The authors show that matching guided velocities directly is under-identified at the branch level, since positive- and negative-branch errors can offset each other.",
					"summary_zh": "on-policy 蒸馏沿学生自身生成的轨迹去查询教师模型，但它在无分类器引导下的行为一直缺乏理解。作者指出，直接匹配引导后的速度在分支层面是欠定的，因为正分支与负分支的误差可以相互抵消。",
					"summary_ja": "オンポリシー蒸留は学生自身の軌跡に沿って教師に問い合わせて拡散モデルを適応させるが、classifier-free guidance の下での挙動は十分に理解されていない。著者らは、ガイド後の速度を直接一致させる目的関数がブランチ水準では劣決定であり、正・負ブランチの誤差が互いに打ち消し合い得ることを示す。"
				},
				{
					"rank": 5,
					"title": "Progress Reward Modeling for Robotic Learning: A Comprehensive Survey",
					"title_zh": "机器人学习中的进度奖励建模：一份综述",
					"title_ja": "ロボット学習における進捗報酬モデリング: 包括的サーベイ",
					"url": "https://huggingface.co/papers/2607.21655",
					"arxiv_id": "2607.21655",
					"upvotes": 55,
					"first_author": "Jianshu Zhang",
					"org": "Northwestern University",
					"summary": "A survey of reward models that score progress rather than completion. A terminal success signal only says whether a task finished, not whether behaviour is advancing, stalling, or undoing earlier progress. The authors note the literature still lacks a shared framework, with methods differing in observations and goal specification.",
					"summary_zh": "一篇关于「衡量进度而非完成度」的奖励模型综述。终止性成功信号只能说明任务是否完成，无法说明行为是在推进、停滞，还是在抹掉此前的进展。作者指出该领域仍缺乏统一框架，各方法在观测与目标设定上各不相同。",
					"summary_ja": "完了ではなく進捗を評価する報酬モデルのサーベイ。終端の成功信号はタスクが終わったかしか示さず、行動が前進しているのか、停滞しているのか、以前の進捗を打ち消しているのかを語らない。著者らは、観測や目標指定が手法ごとに異なり、共通の枠組みがまだないと指摘する。"
				}
			],
			"updated_at": "2026-07-28 12:40 PDT"
		},
		{
			"date": "2026-07-27",
			"items": [
				{
					"rank": 1,
					"title": "Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills",
					"title_zh": "Skill Self-Play：用协同进化的技能推进大模型能力边界",
					"title_ja": "Skill Self-Play: 共進化するスキルで LLM の能力限界を押し広げる",
					"url": "https://huggingface.co/papers/2607.22529",
					"arxiv_id": "2607.22529",
					"upvotes": 16,
					"first_author": "Siyuan Huang",
					"org": "QwenBusinessUnit-Edu",
					"summary": "The paper targets the dilemma in self-evolving LLM training: environment-bound methods get precise feedback but stay narrow, while open-ended self-generation broadens tasks and loses reliable verification. It proposes agent skills as the middle ground, since each skill gives deep and verifiable execution in its own domain.",
					"summary_zh": "论文针对大模型自我进化训练中的两难：绑定环境的方法反馈精确但领域狭窄，开放式自生成任务面广却缺乏可靠验证。作者提出以 agent skills 作为折中，因为每个技能都能在自己的领域内提供可深入验证的执行过程。",
					"summary_ja": "自己進化型 LLM 学習のジレンマを扱う。環境に紐づく手法は正確なフィードバックを得られるが領域が狭く、オープンな自己生成はタスクの幅は広がるが検証の信頼性を失う。著者はその中間解としてエージェントスキルを提案する。各スキルが自領域内で深く検証可能な実行を与えるためである。"
				},
				{
					"rank": 2,
					"title": "DataPrep-Bench: Benchmarking LLMs as Training Data Preparators",
					"title_zh": "DataPrep-Bench：评测大模型作为训练数据准备者的能力",
					"title_ja": "DataPrep-Bench: 学習データの準備者としての LLM を評価する",
					"url": "https://huggingface.co/papers/2607.20465",
					"arxiv_id": "2607.20465",
					"upvotes": 14,
					"first_author": "Hao Liang",
					"summary": "A benchmark for how well LLMs, agents, and data-centric workflows prepare training data end to end. It splits the job into data construction, which turns raw sources into supervised data, and data quality evaluation, which predicts a candidate dataset's downstream training value before training runs.",
					"summary_zh": "一个端到端评测基准，衡量大模型、智能体与以数据为中心的流程准备训练数据的能力。它把这项工作拆成两部分：把原始素材转成监督数据的数据构建，以及在训练前预测候选数据集下游训练价值的数据质量评估。",
					"summary_ja": "LLM、エージェント、データ中心のワークフローが学習データをどれだけ端から端まで準備できるかを測るベンチマーク。作業を、生データを教師データに変換するデータ構築と、学習前に候補データセットの下流での価値を予測するデータ品質評価の 2 つに分ける。"
				},
				{
					"rank": 3,
					"title": "Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning",
					"title_zh": "Molt：面向智能体强化学习的可扩展 PyTorch 原生训练框架",
					"title_ja": "Molt: エージェント型強化学習のためのスケーラブルな PyTorch ネイティブ学習フレームワーク",
					"url": "https://huggingface.co/papers/2607.21653",
					"arxiv_id": "2607.21653",
					"upvotes": 10,
					"first_author": "Jian Hu",
					"org": "NVIDIA",
					"summary": "NVIDIA researchers argue that agentic RL work means constant algorithm changes, and that mainstream frameworks make each change thread through trainer, distributed backend, and rollout glue. Molt keeps the codebase small enough for a researcher, or an AI coding assistant, to hold in full and trace end to end.",
					"summary_zh": "英伟达的研究者指出，智能体强化学习的工作意味着算法不断改动，而主流框架让每次改动都要穿过训练器、分布式后端和 rollout 胶水层。Molt 把代码库压到研究者（或 AI 编程助手）能完整装进脑子、端到端追踪的规模。",
					"summary_ja": "NVIDIA の研究者は、エージェント型 RL の研究は絶え間ない アルゴリズム変更であり、主流フレームワークではその変更がトレーナー、分散バックエンド、ロールアウトの繋ぎ込みを貫いてしまうと指摘する。Molt は研究者や AI コーディングアシスタントが全体を把握し端から端まで追える規模にコードベースを抑える。"
				},
				{
					"rank": 4,
					"title": "Scaling Native Multimodal Pre-Training From Scratch",
					"title_zh": "从零开始扩展原生多模态预训练",
					"title_ja": "ネイティブなマルチモーダル事前学習をゼロからスケールさせる",
					"url": "https://huggingface.co/papers/2607.22043",
					"arxiv_id": "2607.22043",
					"upvotes": 7,
					"first_author": "Haoyuan Wu",
					"org": "Tencent Hunyuan",
					"summary": "Native multimodal pre-training trains from scratch on multimodal inputs, avoiding the text-only limits and late-fusion asymmetries of standard LLMs. The scaling properties of that paradigm are still uncharacterized, so this work investigates the optimal model size and token allocation for it.",
					"summary_zh": "原生多模态预训练直接从零在多模态输入上训练，避开纯文本训练的局限与后融合架构固有的优化不对称。该范式的扩展规律此前缺乏系统刻画，本文据此研究其最优模型规模与 token 分配。",
					"summary_ja": "ネイティブなマルチモーダル事前学習は、マルチモーダル入力でゼロから学習し、テキストのみの制約や後期融合構造の最適化の非対称性を回避する。このパラダイムのスケーリング特性は未解明であり、本研究は最適なモデル規模とトークン配分を調べる。"
				},
				{
					"rank": 5,
					"title": "LAMAR: An Open Language-Aware Multilingual Alignment Reranker",
					"title_zh": "LAMAR：一个开放的语言感知多语种对齐重排器",
					"title_ja": "LAMAR: 言語を考慮したオープンな多言語アラインメント・リランカー",
					"url": "https://huggingface.co/papers/2607.22042",
					"arxiv_id": "2607.22042",
					"upvotes": 6,
					"first_author": "Seongtae Hong",
					"org": "NLP & AI - Korea University",
					"summary": "In multilingual RAG a retriever returns documents in several languages, which are then reranked. The authors show existing rerankers do not consistently prefer documents in the query's own language when semantically equivalent ones exist, even though document language affects the generated answer, and they release LAMAR in response.",
					"summary_zh": "在多语种 RAG 中，检索器会返回多种语言的文档，随后交由重排器排序。作者发现，在存在语义等价文档时，现有重排器并不稳定地优先选择与查询同语种的文档，尽管文档语言会影响最终生成的答案；为此他们发布了 LAMAR。",
					"summary_ja": "多言語 RAG では検索器が複数言語の文書を返し、その後リランクされる。著者らは、意味的に等価な文書がある場合でも既存のリランカーがクエリと同じ言語の文書を一貫して優先しないことを示す。文書の言語は生成される回答に影響するため、対応として LAMAR を公開する。"
				}
			],
			"updated_at": "2026-07-27 00:40 PDT"
		},
		{
			"date": "2026-07-25",
			"items": [
				{
					"rank": 1,
					"title": "NVIDIA-labs OO Agents: Native Python Object-Oriented Agents",
					"url": "https://huggingface.co/papers/2607.20709",
					"arxiv_id": "2607.20709",
					"upvotes": 19,
					"first_author": "Paul Furgale",
					"org": "NVIDIA",
					"summary": "Nvidia proposes NOOA, a framework in which an agent is simply a Python object: methods are its actions, fields its state, docstrings its prompts and type annotations its contracts. A method whose body is only an ellipsis gets completed at runtime by an LLM loop, while normal method bodies stay deterministic Python. The aim is to collapse prompt templates, tool schemas and workflow graphs into one artifact.",
					"summary_zh": "英伟达提出 NOOA 框架：一个智能体就是一个 Python 对象，方法即动作，字段即状态，文档字符串即提示词，类型标注即契约。方法体若只写省略号，就在运行时交由大模型循环补全；写了正常实现的方法则仍是确定性的 Python 代码。其目标是把提示模板、工具 schema 和工作流图收拢成同一份产物。",
					"summary_ja": "エヌビディアは、エージェントを単なる Python オブジェクトとして表現するフレームワーク NOOA を提案した。メソッドが行動、フィールドが状態、docstring がプロンプト、型注釈が契約にあたる。本体が省略記号だけのメソッドは実行時に LLM ループが補完し、通常の実装を持つメソッドは決定的な Python のまま残る。プロンプトテンプレートやツールスキーマ、ワークフローグラフを一つの成果物にまとめることを狙う。",
					"title_zh": "NVIDIA OO Agents：以原生 Python 对象构建智能体",
					"title_ja": "NVIDIA OO Agents: ネイティブ Python のオブジェクト指向エージェント"
				},
				{
					"rank": 2,
					"title": "Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction",
					"url": "https://huggingface.co/papers/2607.20911",
					"arxiv_id": "2607.20911",
					"upvotes": 19,
					"first_author": "Tencent WorkBuddy Bench Team",
					"org": "Tencent",
					"summary": "Tencent released an evaluation suite for coding agents spanning four work domains: code, web, office and security. Rather than reusing public issue text, each task is reverse-engineered from a real commit, pull request or business scenario and rewritten as a short colloquial request, so the prompt cannot be recovered from training data. The report also gives a cross-model leaderboard.",
					"summary_zh": "腾讯发布面向编码智能体的评测套件，覆盖代码、Web、办公与安全四个工作领域。任务并非取自公开的 issue 文本，而是从真实的提交、合并请求或业务场景反向构造，再改写成简短口语化的请求，使提示词无法从训练数据中还原。报告同时给出了构建方法、评分协议和跨模型排行榜。",
					"summary_ja": "テンセントは、コード・Web・オフィス・セキュリティの 4 領域にまたがるコーディングエージェント評価スイートを公開した。公開 issue の文面を流用せず、実際のコミットやプルリクエスト、業務シナリオから逆算してタスクを構成し、短い口語調の依頼文に書き換えることで、プロンプトが学習データから復元できないようにしている。報告書には構築手法、採点プロトコル、モデル横断のリーダーボードがまとめられている。",
					"title_zh": "腾讯 WorkBuddy Bench：抗污染任务构造的多领域编码智能体基准",
					"title_ja": "Tencent WorkBuddy Bench: 汚染耐性のあるタスク構成による多領域コーディングエージェント評価"
				},
				{
					"rank": 3,
					"title": "Color Pass-Through via Camera-Display Coupling",
					"url": "https://huggingface.co/papers/2607.12746",
					"arxiv_id": "2607.12746",
					"upvotes": 18,
					"first_author": "Ruikang Li",
					"org": "MMLab-CUHK",
					"summary": "A scene shot on a phone and shown on the same phone's screen still looks noticeably off in color, brightness and contrast. The authors trace the gap to pipelines that calibrate the camera and the display separately and then bridge them with low-dimensional color transforms, which creates information bottlenecks and compounding error. They propose treating capture and display as one coupled system instead.",
					"summary_zh": "同一部手机拍下的画面，在自己的屏幕上显示时，色彩、亮度和对比度往往与实景明显不同。作者指出问题出在现有流程：相机与显示分别标定，再用低维色彩变换连接，由此产生信息瓶颈和误差累积。他们主张把拍摄与显示当作一个耦合系统统一处理。",
					"summary_ja": "同じスマートフォンで撮影した情景を自機の画面で見ると、色や明るさ、コントラストが実景と目に見えて食い違う。著者らは、カメラとディスプレイを別々に校正し、低次元の色変換でつなぐ既存パイプラインが情報のボトルネックと誤差の累積を生んでいると指摘する。そこで撮影と表示を一つの結合系として扱う手法を提案している。",
					"title_zh": "相机与显示耦合的色彩直通",
					"title_ja": "カメラとディスプレイの結合による色のパススルー"
				},
				{
					"rank": 4,
					"title": "SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation",
					"url": "https://huggingface.co/papers/2607.21553",
					"arxiv_id": "2607.21553",
					"upvotes": 18,
					"first_author": "Junsong Chen",
					"org": "NVIDIA",
					"summary": "SANA-Video 2.0 is a hybrid video diffusion transformer at 5B and 14B scale that generates up to 720p video on a single GPU. It mixes gated linear attention for the bulk of token mixing with periodic gated-softmax anchors at a 3:1 ratio, restoring the full-rank token interactions pure linear attention loses. The claim is full-softmax quality with linear attention's long-sequence scaling.",
					"summary_zh": "SANA-Video 2.0 是一个混合式视频扩散 Transformer，提供 5B 与 14B 两种规模，可在单张 GPU 上生成最高 720p 的视频。它以门控线性注意力承担大部分 token 混合，并按 3:1 的比例周期性插入门控 softmax 锚点，补回纯线性注意力所缺失的满秩 token 交互。作者称其画质可比肩全 softmax 模型，同时保留线性注意力在长序列上的扩展优势。",
					"summary_ja": "SANA-Video 2.0 は 5B と 14B の規模を持つハイブリッド動画拡散トランスフォーマーで、単一 GPU で最大 720p の動画を生成する。トークン混合の大半をゲート付き線形注意が担い、3:1 の比率でゲート付き softmax のアンカーを周期的に挟むことで、純粋な線形注意では失われる全ランクのトークン相互作用を回復させる。フル softmax 並みの品質と、線形注意の長系列スケーリングの両立を主張している。",
					"title_zh": "SANA-Video 2.0：用带注意力残差的混合线性注意力实现高效视频生成",
					"title_ja": "SANA-Video 2.0: 注意残差を備えたハイブリッド線形注意による効率的な動画生成"
				},
				{
					"rank": 5,
					"title": "LLMs Get Lost in Evolving User Intent",
					"url": "https://huggingface.co/papers/2607.20734",
					"arxiv_id": "2607.20734",
					"upvotes": 17,
					"first_author": "Jihoon Tack",
					"org": "Microsoft Research",
					"summary": "Real users rarely state what they want upfront; they disclose, revise and reshape it as a conversation unfolds. Yet models are still mostly evaluated and trained on single-turn, fully-specified prompts. Microsoft Research introduces a framework that turns fully-specified tasks into evolving ones to test how well models track intent that keeps moving.",
					"summary_zh": "真实用户很少一开始就把需求说清楚，而是在对话过程中逐步透露、修改乃至重塑意图。但当前模型的评测与训练，仍以单轮、需求完整给定的提示为主。微软研究院提出一套框架，把完整给定的任务改造成随对话演化的任务，用以检验模型能否跟上不断变化的用户意图。",
					"summary_ja": "現実のユーザーは最初に要件を言い切ることはまれで、会話の中で少しずつ明かし、修正し、作り変えていく。それにもかかわらずモデルの評価と学習は、いまなお単一ターンで要件が完全に指定されたプロンプトが中心である。マイクロソフトリサーチは、完全指定のタスクを変化するタスクへ変換する枠組みを導入し、動き続ける意図をモデルがどこまで追えるかを検証した。",
					"title_zh": "用户意图不断变化时，大模型会迷失方向",
					"title_ja": "LLM は移り変わるユーザーの意図を見失う"
				}
			],
			"updated_at": "2026-07-25 11:23 PDT"
		},
		{
			"date": "2026-07-24",
			"items": [
				{
					"rank": 1,
					"title": "AREX: Towards a Recursively Self-Improving Agent for Deep Research",
					"title_zh": "AREX:面向深度研究的递归自我改进 agent",
					"title_ja": "AREX:ディープリサーチのための再帰的自己改善エージェント",
					"url": "https://huggingface.co/papers/2607.21461",
					"arxiv_id": "2607.21461",
					"upvotes": 115,
					"first_author": "Shuqi Lu",
					"org": "Beijing Academy of Artificial Intelligence",
					"summary": "AREX proposes an agent for deep research that recursively improves itself, refining its own research strategies over successive iterations rather than staying fixed. It drew the day's most upvotes on Hugging Face by a wide margin. The work is from the Beijing Academy of Artificial Intelligence.",
					"summary_zh": "AREX 提出了一种面向深度研究的 agent,它能递归地自我改进,在一次次迭代中优化自身的研究策略,而非固定不变。它以明显优势获得当日 Hugging Face 最多的点赞。该研究来自北京智源人工智能研究院(BAAI)。",
					"summary_ja": "AREXは、ディープリサーチのためのエージェントを提案する。固定されたままではなく、反復のたびに自らの研究戦略を洗練させ、再帰的に自己改善する。当日のHugging Faceで大差の最多支持を集めた。北京智源人工知能研究院(BAAI)による研究だ。"
				},
				{
					"rank": 2,
					"title": "ReferTrack: Referring Then Tracking for Embodied Visual Tracking",
					"title_zh": "ReferTrack:先指代后跟踪的具身视觉追踪",
					"title_ja": "ReferTrack:参照してから追跡する身体化視覚トラッキング",
					"url": "https://huggingface.co/papers/2607.20061",
					"arxiv_id": "2607.20061",
					"upvotes": 43,
					"first_author": "Hanjing Ye",
					"org": "Tencent",
					"summary": "ReferTrack tackles embodied visual tracking by first resolving a natural-language reference to the target, then tracking it, rather than tracking blindly. The two-stage design aims to make an embodied agent follow the object a user actually means. The paper is from Tencent.",
					"summary_zh": "ReferTrack 处理具身视觉追踪:先根据自然语言指代确定目标,再进行跟踪,而非盲目追踪。这种两阶段设计意在让具身 agent 跟随用户真正所指的物体。论文来自腾讯。",
					"summary_ja": "ReferTrackは身体化された視覚トラッキングに取り組む。盲目的に追跡するのではなく、まず自然言語の参照から対象を特定し、その上で追跡する。この二段階の設計は、身体化エージェントがユーザーの意図する物体を追えるようにすることを狙う。テンセントによる論文だ。"
				},
				{
					"rank": 3,
					"title": "K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs",
					"title_zh": "K12-KGraph:用于评测与训练教育类 LLM 的课程对齐知识图谱",
					"title_ja": "K12-KGraph:教育向けLLMの評価と訓練のためのカリキュラム整合ナレッジグラフ",
					"url": "https://huggingface.co/papers/2605.09635",
					"arxiv_id": "2605.09635",
					"upvotes": 40,
					"first_author": "Hao Liang",
					"org": "Peking University",
					"summary": "K12-KGraph introduces a knowledge graph aligned to K-12 curricula, built to benchmark and train educational large language models against structured grade-level content. It targets the gap between general LLM knowledge and what students are actually taught. The work comes from Peking University.",
					"summary_zh": "K12-KGraph 提出了一个与 K-12 课程对齐的知识图谱,用于依据结构化的年级内容来评测和训练教育类大语言模型。它针对的是通用 LLM 知识与学生实际所学之间的差距。该工作来自北京大学。",
					"summary_ja": "K12-KGraphは、K12カリキュラムに整合したナレッジグラフを提案し、構造化された学年別コンテンツに照らして教育向け大規模言語モデルを評価・訓練する。汎用LLMの知識と、生徒が実際に学ぶ内容との差を埋めることを狙う。北京大学による研究だ。"
				},
				{
					"rank": 4,
					"title": "Visual Contrastive Self-Distillation",
					"title_zh": "视觉对比自蒸馏",
					"title_ja": "視覚的コントラスティブ自己蒸留",
					"url": "https://huggingface.co/papers/2607.21556",
					"arxiv_id": "2607.21556",
					"upvotes": 39,
					"first_author": "Yijun Liang",
					"org": "University of Maryland College Park",
					"summary": "Visual Contrastive Self-Distillation presents a self-supervised vision method that combines contrastive learning with self-distillation, letting a model learn visual representations by teaching itself. The approach aims to improve representation quality without extra labels. The paper is from the University of Maryland, College Park.",
					"summary_zh": "《视觉对比自蒸馏》提出一种自监督视觉方法,将对比学习与自蒸馏结合,让模型通过“自我教学”来学习视觉表征。该方法意在无需额外标注即可提升表征质量。论文来自马里兰大学学院市分校。",
					"summary_ja": "「視覚的コントラスティブ自己蒸留」は、コントラスティブ学習と自己蒸留を組み合わせた自己教師あり視覚手法を提案し、モデルが自らを教えることで視覚表現を学習できるようにする。追加のラベルなしに表現の質を高めることを狙う。メリーランド大学カレッジパーク校による論文だ。"
				},
				{
					"rank": 5,
					"title": "Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text",
					"title_zh": "Show, Don't Tell:在生成像素而非 LLM 文本中评测空间认知",
					"title_ja": "Show, Don't Tell:LLMのテキストではなく生成ピクセルで空間認知を評価する",
					"url": "https://huggingface.co/papers/2607.21072",
					"arxiv_id": "2607.21072",
					"upvotes": 34,
					"first_author": "Xu Wang",
					"org": "ZJU-OmniAI",
					"summary": "This paper argues spatial cognition should be evaluated by what generative models draw, not what LLMs say, and builds a benchmark that scores spatial understanding directly in generated pixels. It contends text-based tests miss capabilities and failures visible only in imagery. The work is from ZJU-OmniAI.",
					"summary_zh": "该论文主张:空间认知应通过生成模型“画出”的内容来评测,而非 LLM“说出”的内容,并构建了一个直接在生成像素中评分空间理解的基准。作者认为,基于文本的测试会遗漏只有在图像中才可见的能力与缺陷。该工作来自 ZJU-OmniAI。",
					"summary_ja": "本論文は、空間認知はLLMが「語る」内容ではなく生成モデルが「描く」内容で評価すべきだと主張し、生成ピクセルの中で空間理解を直接採点するベンチマークを構築する。テキストベースのテストは、画像でしか見えない能力や失敗を見落とすと論じる。ZJU-OmniAIによる研究だ。"
				}
			],
			"updated_at": "2026-07-24 15:44 PDT"
		},
		{
			"date": "2026-07-23",
			"items": [
				{
					"rank": 1,
					"title": "Generative World Renderer at the Speed of Play",
					"title_zh": "以游玩速度运行的生成式世界渲染器",
					"title_ja": "プレイ速度で動く生成的ワールドレンダラー",
					"url": "https://huggingface.co/papers/2607.18703",
					"arxiv_id": "2607.18703",
					"upvotes": 67,
					"first_author": "Guixu Lin",
					"org": "Alaya Lab",
					"summary": "The paper introduces a generative world renderer that produces game-like visual worlds fast enough to run at interactive play speed. It frames real-time world generation as a rendering problem a single model can drive rather than a slow offline pipeline. The work drew the day's most upvotes on Hugging Face.",
					"summary_zh": "论文提出了一种生成式世界渲染器,能生成类似游戏的视觉世界,速度快到可在交互式游玩帧率下运行。它把实时世界生成视为可由单一模型驱动的渲染问题,而非缓慢的离线流程。该研究获得了当日 Hugging Face 上最多的点赞。",
					"summary_ja": "本論文は、ゲームのような視覚世界を対話的なプレイ速度で生成できる生成的ワールドレンダラーを提案する。リアルタイムの世界生成を、遅いオフライン処理ではなく単一モデルで駆動できるレンダリング問題として定式化する。当日Hugging Faceで最多の支持を集めた。"
				},
				{
					"rank": 2,
					"title": "Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing",
					"title_zh": "Mage-Flow:面向图像生成与编辑的高效原生分辨率基础模型",
					"title_ja": "Mage-Flow:画像生成・編集のための効率的なネイティブ解像度基盤モデル",
					"url": "https://huggingface.co/papers/2607.19064",
					"arxiv_id": "2607.19064",
					"upvotes": 60,
					"first_author": "Xinjie Zhang",
					"org": "Microsoft",
					"summary": "Microsoft researchers present Mage-Flow, a compact foundation model for image generation and instruction-based editing that operates at native resolution instead of upscaling from a fixed low-resolution latent. The paper emphasizes efficiency at full resolution over raw parameter count. It was among the day's most-upvoted submissions.",
					"summary_zh": "微软研究者提出 Mage-Flow,一款用于图像生成和基于指令编辑的紧凑基础模型,直接在原生分辨率上工作,而非从固定的低分辨率隐空间上采样。论文强调全分辨率下的效率,而非单纯堆参数。它是当日点赞最多的投稿之一。",
					"summary_ja": "マイクロソフトの研究者らは、固定の低解像度潜在からのアップスケールではなくネイティブ解像度で動作する、画像生成と指示ベース編集向けのコンパクトな基盤モデルMage-Flowを提案した。論文はパラメータ数よりも全解像度での効率を重視する。当日最も支持された投稿の一つだった。"
				},
				{
					"rank": 3,
					"title": "SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD",
					"title_zh": "SLAI T-Rex:在昇腾 SuperPOD 上对 DeepSeek-V4 系列做全参数后训练",
					"title_ja": "SLAI T-Rex:Ascend SuperPODでのDeepSeek-V4ファミリーの全パラメータ事後学習",
					"url": "https://huggingface.co/papers/2607.20145",
					"arxiv_id": "2607.20145",
					"upvotes": 54,
					"first_author": "Dongfang Li",
					"summary": "The paper documents full-parameter post-training of the DeepSeek-V4 model family on Huawei's Ascend SuperPOD, a large-scale run outside the usual Nvidia stack. It details the systems engineering needed to fine-tune frontier-scale models on Ascend hardware. It ranked among the day's top papers.",
					"summary_zh": "论文记录了在华为昇腾 SuperPOD 上对 DeepSeek-V4 模型系列进行全参数后训练的过程,这是一次脱离常见 Nvidia 生态的大规模实践。它详述了在昇腾硬件上微调前沿规模模型所需的系统工程。该研究位列当日热门论文之一。",
					"summary_ja": "本論文は、通常のNvidiaスタック外での大規模な取り組みとして、HuaweiのAscend SuperPOD上でDeepSeek-V4ファミリーの全パラメータ事後学習を行った過程を記録する。Ascendハードウェアでフロンティア規模のモデルを微調整するために必要なシステム工学を詳述する。当日の上位論文に入った。"
				},
				{
					"rank": 4,
					"title": "Subliminal Clocks: Latent Time Modelling in Diffusion Language Models",
					"title_zh": "潜时钟:扩散语言模型中的隐式时间建模",
					"title_ja": "サブリミナル・クロック:拡散言語モデルにおける潜在的な時間モデリング",
					"url": "https://huggingface.co/papers/2607.01774",
					"arxiv_id": "2607.01774",
					"upvotes": 36,
					"first_author": "Maximo Eduardo Rulli",
					"org": "Sapienza University of Rome",
					"summary": "This paper studies how diffusion language models implicitly track progress through generation, proposing latent 'subliminal clocks' that encode timing inside the denoising process. It offers a lens on how such models order their output without explicit autoregression. The work came from Sapienza University of Rome.",
					"summary_zh": "论文研究扩散语言模型如何隐式地追踪生成进度,提出在去噪过程内部编码时序的潜在“潜时钟”。它为理解这类模型在没有显式自回归的情况下如何安排输出顺序提供了视角。该工作来自罗马大学(Sapienza)。",
					"summary_ja": "本論文は、拡散言語モデルが生成の進行を暗黙的にどう追跡するかを調べ、ノイズ除去過程の内部に時間を符号化する潜在的な「サブリミナル・クロック」を提案する。明示的な自己回帰なしにこれらのモデルが出力を順序付ける仕組みへの視点を与える。ローマ・サピエンツァ大学による研究だ。"
				},
				{
					"rank": 5,
					"title": "Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning",
					"title_zh": "陈旧亦稳定:用陈旧度自适应信赖域稳定异步强化学习",
					"title_ja": "古くても安定:非同期強化学習を安定化する陳腐化適応型トラスト領域",
					"url": "https://huggingface.co/papers/2607.18722",
					"arxiv_id": "2607.18722",
					"upvotes": 31,
					"first_author": "Junyao Yang",
					"org": "Tencent Hunyuan",
					"summary": "Tencent Hunyuan researchers propose a staleness-adaptive trust region to stabilize asynchronous reinforcement learning, where updates computed on outdated policy copies can destabilize training. The method scales each update's trust region by how stale it is. It aims at faster yet more stable large-scale RL.",
					"summary_zh": "腾讯混元的研究者提出一种陈旧度自适应信赖域方法,用以稳定异步强化学习——在异步训练中,基于过时策略副本计算的更新可能破坏训练稳定性。该方法按每次更新的陈旧程度来缩放其信赖域,目标是让大规模强化学习既更快又更稳。",
					"summary_ja": "テンセント混元の研究者らは、非同期強化学習を安定化する陳腐化適応型トラスト領域を提案する。非同期学習では、古いポリシーのコピーで計算された更新が学習を不安定にしうる。本手法は各更新の陳腐度に応じてトラスト領域を調整する。大規模RLをより速く、かつ安定にすることを狙う。"
				}
			],
			"updated_at": "2026-07-23 23:34 PDT"
		},
		{
			"date": "2026-07-22",
			"items": [
				{
					"rank": 1,
					"title": "RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model",
					"url": "https://huggingface.co/papers/2607.17977",
					"upvotes": 187,
					"arxiv_id": "2607.17977",
					"summary": "DAMO Academy presents RynnBrain 1.1, a family of embodied foundation models spanning 2B to 122B parameters, aimed at more capable and generalizable control for physical agents.",
					"summary_zh": "达摩院提出 RynnBrain 1.1,一族参数量从 2B 到 122B 的具身基础模型,目标是让物理智能体的控制更强、更具泛化性。",
					"summary_ja": "DAMOアカデミーは、2Bから122Bパラメータに及ぶ身体性基盤モデル群RynnBrain 1.1を発表。物理エージェントの制御をより高性能かつ汎用的にすることを狙う。",
					"title_zh": "RynnBrain 1.1:迈向更强、更通用的具身基础模型",
					"title_ja": "RynnBrain 1.1:より高性能で汎用的な身体性基盤モデルへ",
					"first_author": "Kehan Li",
					"org": "DAMO Academy"
				},
				{
					"rank": 2,
					"title": "ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU",
					"url": "https://huggingface.co/papers/2607.19191",
					"upvotes": 170,
					"arxiv_id": "2607.19191",
					"summary": "Alibaba's AMAP CV Lab introduces ABot-World-0, an action-conditioned video world model for real-time, long-horizon interactive rollout - a step toward world models robots can plan against.",
					"summary_zh": "阿里 AMAP CV 实验室提出 ABot-World-0,一个以动作为条件的视频世界模型,支持实时、长时程的交互式推演——朝着机器人可据以规划的世界模型迈进一步。",
					"summary_ja": "アリババのAMAP CV Labは、行動条件付きの動画世界モデルABot-World-0を発表。リアルタイムかつ長期の対話的ロールアウトを可能にし、ロボットが計画に使える世界モデルへ一歩近づく。",
					"title_zh": "ABot-World-0:面向机器人的无限交互式世界模型",
					"title_ja": "ABot-World-0:ロボット向け無限インタラクティブ世界モデル",
					"first_author": "Fan Jiang",
					"org": "Alibaba AMAP CV Lab"
				},
				{
					"rank": 3,
					"title": "DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines",
					"url": "https://huggingface.co/papers/2607.16617",
					"upvotes": 122,
					"arxiv_id": "2607.16617",
					"summary": "From Peking University, DataFlow-Harness is a grounded code-agent platform for automating data-processing workflows, giving agents an executable, verifiable substrate instead of free-form code.",
					"summary_zh": "北京大学提出 DataFlow-Harness,一个可落地的代码智能体平台,用于自动化数据处理流程,为智能体提供可执行、可验证的底座,而非自由生成的代码。",
					"summary_ja": "北京大学によるDataFlow-Harnessは、データ処理ワークフローを自動化する実地的なコードエージェント基盤で、自由記述のコードではなく実行・検証可能な土台をエージェントに与える。",
					"title_zh": "DataFlow-Harness:面向数据处理的可落地代码智能体平台",
					"title_ja": "DataFlow-Harness:データ処理向けの実地コードエージェント基盤",
					"first_author": "Runming He",
					"org": "Peking University"
				},
				{
					"rank": 4,
					"title": "SWE-Pruner Pro: The Coder LLM Already Knows What to Prune",
					"url": "https://huggingface.co/papers/2607.18213",
					"upvotes": 71,
					"arxiv_id": "2607.18213",
					"summary": "ByteDance's SWE-Pruner Pro prunes long context for coding agents by using the coder model's own signals about what matters, cutting tokens while keeping the context an agent actually needs.",
					"summary_zh": "字节跳动的 SWE-Pruner Pro 借助编码模型自身对\"什么重要\"的信号来裁剪长上下文,在削减 token 的同时保留智能体真正需要的上下文。",
					"summary_ja": "バイトダンスのSWE-Pruner Proは、何が重要かというコーダーモデル自身のシグナルを使ってコーディングエージェントの長い文脈を刈り込み、トークンを削りつつ本当に必要な文脈を残す。",
					"title_zh": "SWE-Pruner Pro:编码智能体其实已知道该保留什么",
					"title_ja": "SWE-Pruner Pro:コーディングエージェントは残すべき文脈を既に知っている",
					"first_author": "Yuhang Wang",
					"org": "ByteDance"
				},
				{
					"rank": 5,
					"title": "Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers",
					"url": "https://huggingface.co/papers/2607.19139",
					"upvotes": 67,
					"arxiv_id": "2607.19139",
					"summary": "This paper studies text-to-image diffusion transformers and argues the template tokens they process act as implicit semantic registers, offering a new handle on how DiTs bind text to image.",
					"summary_zh": "该论文研究文生图扩散 Transformer,提出其处理的模板 token 充当隐式的语义寄存器,为理解 DiT 如何把文本绑定到图像提供了新的抓手。",
					"summary_ja": "本論文はテキスト画像拡散トランスフォーマー(DiT)を調べ、処理されるテンプレートトークンが暗黙の意味レジスタとして働くと論じ、DiTがテキストを画像へ結びつける仕組みへの新たな手がかりを与える。",
					"title_zh": "文本模板 token 是隐式的语义寄存器",
					"title_ja": "テキストテンプレートトークンは暗黙の意味レジスタである",
					"first_author": "Maohua Li",
					"org": "RTP-LLM"
				}
			]
		},
		{
			"date": "2026-07-21",
			"items": [
				{
					"rank": 1,
					"title": "TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs",
					"url": "https://huggingface.co/papers/2607.17423",
					"upvotes": 137,
					"arxiv_id": "2607.17423",
					"summary": "Video MLLMs can describe what happens in a video but rarely when the supporting evidence occurs. TimeLens2 targets generalist temporal grounding - locating the moments that back an answer across video tasks.",
					"summary_zh": "视频多模态大模型能描述视频里发生了什么,却很少能说出证据出现在何时。TimeLens2 面向通用时序定位:在各类视频任务中找到支撑答案的具体时刻。",
					"summary_ja": "動画MLLMは何が起きたかは説明できても、その根拠がいつ現れるかはほとんど示せない。TimeLens2は汎用の時間グラウンディング、つまり答えを裏付ける瞬間の特定に挑む。",
					"title_zh": "TimeLens2:面向通用视频时序定位的多模态大模型",
					"title_ja": "TimeLens2:汎用ビデオ時間グラウンディング",
					"first_author": "Yuhan Zhu",
					"org": "Multimedia Computing Group-Nanjing University"
				},
				{
					"rank": 2,
					"title": "RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources",
					"url": "https://huggingface.co/papers/2606.29538",
					"upvotes": 128,
					"arxiv_id": "2606.29538",
					"summary": "From Microsoft Research: distilling executable agent skills from human resources and experience, turning procedures into reusable skill libraries instead of hand-written ones.",
					"summary_zh": "微软研究院出品:从人类资源与经验中蒸馏可执行的智能体技能,把流程沉淀为可复用的技能库,而非手写技能。",
					"summary_ja": "Microsoft Researchによる研究。人間の資産や経験から実行可能なエージェントスキルを蒸留し、手書きではなく再利用可能なスキルライブラリへ手順を昇華する。",
					"title_zh": "RESOURCE2SKILL:从人类资源中蒸馏可执行的智能体技能",
					"title_ja": "RESOURCE2SKILL:人間の資産から実行可能なエージェントスキルを蒸留",
					"first_author": "Yijia Fan",
					"org": "Microsoft Research"
				},
				{
					"rank": 3,
					"title": "RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM",
					"url": "https://huggingface.co/papers/2607.11683",
					"upvotes": 128,
					"arxiv_id": "2607.11683",
					"summary": "A multi-step GraphRAG engine built around a compact, domain-adaptive knowledge graph, addressing the cost and rigidity of existing graph-construction pipelines.",
					"summary_zh": "围绕紧凑、领域自适应知识图谱构建的多步 GraphRAG 引擎,针对现有图构建管线成本高、灵活性差的问题。",
					"summary_ja": "コンパクトでドメイン適応型のナレッジグラフを核とする多段GraphRAGエンジン。既存のグラフ構築パイプラインのコストと硬直性に対処する。",
					"title_zh": "RAGU:带紧凑领域自适应知识图谱的多步 GraphRAG 引擎",
					"title_ja": "RAGU:コンパクトなドメイン適応グラフを備えた多段GraphRAGエンジン",
					"first_author": "Mikhail Komarov",
					"org": "Novosibirsk State University"
				},
				{
					"rank": 4,
					"title": "EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World",
					"url": "https://huggingface.co/papers/2607.17250",
					"upvotes": 71,
					"arxiv_id": "2607.17250",
					"summary": "An open-schema framework and benchmark for character-and-world co-evolution in interactive literary worlds, from Tencent. Characters and their fictional settings change each other over time.",
					"summary_zh": "腾讯提出的开放模式框架与基准:交互式文学世界中的角色与世界共同演化,人物与虚构世界随时间互相塑造。",
					"summary_ja": "テンセントによるオープンスキーマの枠組みとベンチマーク。インタラクティブな物語世界でキャラクターと世界が時間とともに相互に形づくる共進化を扱う。",
					"title_zh": "EvolvingWorld:角色与世界共同演化的开放模式框架",
					"title_ja": "EvolvingWorld:キャラクターと世界が共進化するオープンスキーマ枠組み",
					"first_author": "Qing Zong",
					"org": "Tencent"
				},
				{
					"rank": 5,
					"title": "DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment",
					"url": "https://huggingface.co/papers/2607.07820",
					"upvotes": 69,
					"arxiv_id": "2607.07820",
					"summary": "Self-distillation for deep-search agents: instead of fixed teacher-distilled trajectories, agents improve from their own successful search episodes.",
					"summary_zh": "深度搜索智能体的自蒸馏:不依赖固定的教师蒸馏轨迹,智能体从自身成功的搜索过程中持续改进。",
					"summary_ja": "深層検索エージェントの自己蒸留。固定された教師由来の軌跡ではなく、エージェント自身の成功した探索エピソードから学習を重ねる。",
					"title_zh": "DeepSearch-World:深度搜索智能体的自蒸馏训练",
					"title_ja": "DeepSearch-World:深層検索エージェントの自己蒸留",
					"first_author": "Xinyu Geng",
					"org": "HKUST"
				}
			]
		},
		{
			"date": "2026-07-17",
			"items": [
				{
					"rank": 1,
					"title": "VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding",
					"url": "https://huggingface.co/papers/2607.14935",
					"upvotes": 111,
					"arxiv_id": "2607.14935",
					"summary": "A fully open video multimodal LLM covering motion, long-video and streaming understanding, positioned as an open reference stack for video assistants.",
					"first_author": "Xinhao Li",
					"org": "Multimedia Computing Group-Nanjing University",
					"summary_zh": "完全开放的视频多模态大模型，覆盖动作理解、长视频与流式交互，定位为视频助手的开放参考实现。",
					"summary_ja": "動作理解・長尺動画・ストリーミング対話をカバーする完全オープンな動画マルチモーダルLLM。ビデオアシスタントのオープンな参照実装を目指す。",
					"title_zh": "VideoChat3：面向高效通用视频理解的完全开放视频多模态大模型",
					"title_ja": "VideoChat3：効率的で汎用的な動画理解のための完全オープン動画MLLM"
				},
				{
					"rank": 2,
					"title": "LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget",
					"url": "https://huggingface.co/papers/2607.14952",
					"upvotes": 97,
					"arxiv_id": "2607.14952",
					"summary": "Reinforcement-learning post-training beyond 2 million tokens of context under a fixed GPU budget, attacking the widening gap between inference context lengths and what RL training can reach.",
					"first_author": "Changhai Zhou",
					"org": "Mind Lab",
					"summary_zh": "在固定 GPU 预算下把强化学习后训练扩展到 200 万 token 以上的上下文，直面推理上下文长度与 RL 训练能力之间日益扩大的差距。",
					"summary_ja": "固定GPU予算のもとで強化学習ポストトレーニングを200万トークン超のコンテキストへ拡張。推論時コンテキスト長とRL訓練の間で広がるギャップに挑む。",
					"title_zh": "LongStraw：固定 GPU 预算下超 200 万 token 的长上下文强化学习",
					"title_ja": "LongStraw：固定GPU予算で200万トークン超の長文脈RL"
				},
				{
					"rank": 3,
					"title": "SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning",
					"url": "https://huggingface.co/papers/2607.14777",
					"upvotes": 73,
					"arxiv_id": "2607.14777",
					"summary": "Self-evolving on-policy distillation for agentic RL: models trained as interactive agents on long-horizon tasks learn from their own improving policy instead of a fixed teacher.",
					"first_author": "Jinyang Wu",
					"summary_zh": "面向智能体强化学习的自进化在线蒸馏：在长程任务中作为交互式智能体训练的模型，从自身不断改进的策略中学习，而不是依赖固定的教师模型。",
					"summary_ja": "エージェント強化学習のための自己進化型オンポリシー蒸留。長期タスクで対話型エージェントとして訓練されたモデルが、固定の教師ではなく自らの改善する方策から学ぶ。",
					"title_zh": "SEED：面向智能体强化学习的自进化在线蒸馏",
					"title_ja": "SEED：エージェント強化学習のための自己進化型オンポリシー蒸留"
				},
				{
					"rank": 4,
					"title": "SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration",
					"url": "https://huggingface.co/papers/2607.15257",
					"upvotes": 52,
					"arxiv_id": "2607.15257",
					"summary": "Works toward robust open-domain information seeking with tool-integrated LLMs, treating web search as a core model capability rather than a bolted-on tool.",
					"first_author": "Yuyao Zhang",
					"org": "Ant Group",
					"summary_zh": "研究让工具集成的大模型具备稳健的开放域信息检索能力，把网页搜索当作模型的核心能力而不是外挂工具。",
					"summary_ja": "ツール統合LLMによる頑健なオープンドメイン情報探索を目指す研究。ウェブ検索を後付けのツールではなくモデルの中核能力として扱う。",
					"title_zh": "SearchOS-V1：迈向稳健的开放域信息检索",
					"title_ja": "SearchOS-V1：頑健なオープンドメイン情報探索へ"
				},
				{
					"rank": 5,
					"title": "BadWAM: When World-Action Models Dream Right but Act Wrong",
					"url": "https://huggingface.co/papers/2607.15207",
					"upvotes": 37,
					"arxiv_id": "2607.15207",
					"summary": "Examines world-action models that predict the future correctly yet still act wrongly - an embodied-control failure mode where dreaming right does not mean acting right.",
					"first_author": "Qi Li",
					"summary_zh": "研究“世界-行动模型”预测未来正确却行动错误的现象——一种具身控制的失效模式：梦得对不代表做得对。",
					"summary_ja": "世界行動モデル（WAM）が未来を正しく予測しながら誤った行動を取る現象を分析。「正しく夢を見ても正しく動ける」とは限らないという身体性制御の失敗モードを示す。",
					"title_zh": "BadWAM：当世界-行动模型梦得对却做得错",
					"title_ja": "BadWAM：世界行動モデルが正しく夢を見て誤って動くとき"
				}
			]
		}
	]
}
