話題の最新モデル、ブレイクスルー研究、注目ツールの情報をいち早くキャッチ
AnthropicがClaude 4を発表。前世代比で推論能力が45%向上し、画像・動画・音声のマルチモーダル理解が飛躍的に進化。コーディング性能も大幅アップ。
GoogleがGemma 4を公開。400億パラメータ級でありながら量子化技術により一般GPUでも動作。ベンチマークでLlama 4を上回る性能を記録。
As Chinese AI models grow in capability and popularity among U.S. companies, the arguing over what should be done about them has reached a fever pitch.
Synthesia launched AI Roleplay Sessions, an interactive enterprise training platform where employees practice workplace conversations with AI avatars that provide feedback, scoring, and analytics to help companies measure training effectiveness.
OpenAI has come forward to claim responsibility for the Hugging Face breach, saying it was the result of internal testing gone awry.
OpenAI has come forward to claim responsibility for the Hugging Face breach, saying it was the result of internal testing gone awry.
Google released Gemini 3.6 Flash, 3.5 Flash-Lite, and Flash Cyber, but the continued absence of Gemini 3.5 Pro raises fresh questions about its AI strategy.
Treasury Secretary Scott Bessent said the U.S. could sanction Chinese open AI models over alleged IP theft, expanding the Trump administration's campaign to slow China's AI advances.
The final approval settles one case, but it doesn't resolve the broader issue of using copyrighted works to train AI models.
Validating Distributed LLM Serving Benchmarks with NVIDIA srt-slurm, SLURM Recipes, Parameter Sweeps, and Pareto Analysis MarkTechPost
Mortif Technologies' own Large Language Model (LLM) ranked third among open weight models in the glo.. 매일경제
Chinese models are on track to win the agentic AI price war The Strategist | ASPI's analysis and commentary site
Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of reported WER attributable to style mismatch), and unreliable word-level timing. We show that models already encode both styles; the challenge is controlled activation. Using coverage-aware decoder task tokens trained on parallel verbatim/intended transcript pairs, we raise Ge
Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description. This survey focuses on visual humor understanding in single-image and multi-panel artifacts, while treating humor generation as an emerging downstream frontier. We position the literature against prior humor, sarcasm, and general MLLM surveys and organize it using a c
Accurate agricultural field boundary delineation at large scale is a foundational task for food security, supply chain transparency, and carbon accounting. While vision foundation models like SAM show remarkable zero-shot capabilities, they frequently fail in geospatial domains due to topological complexity, cropland texturing patterns, and a lack of physical scale awareness. In this work, we introduce Delineate Anything v2, a globally scalable foundation model designed specifically for wide-are
Optimizer state is the largest single line item in the memory budget of mixture-of-experts (MoE) training: on a 6.78B-parameter MoE language model, AdamW keeps 50.6 GB of first and second moments to update 12.6 GB of bfloat16 weights. We study SkewAdam, an optimizer built on the observation that the three parameter populations of an MoE - the dense backbone, the experts, and the router - differ enough in size and gradient statistics that they should not receive the same state. SkewAdam keeps flo
Alphabet, Google's parent company, is reportedly working on a new chip designed to make its Gemini models run much more efficiently.
Talk of banning Chinese-made open-weight LLMs reveals the challenge of turning AI into a business.
Claude plus a local LLM cuts my AI costs in half, and I'm never going back to cloud-only XDA
Do Large Language Models Think Like Us? Psychology Today
Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Unlike standard diffusion-based approaches, DLMs are not explicitly conditioned on a timestep, raising a natural question: do these models internally represent denoising progress, and how is such information used downstream? In this work, we show that DLMs do in fact encode a latent representation related to the diffusion timestep within their residual streams. We find that this signal can
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or lie about the result. No benchmark measures which of these an uninstructed model chooses. We introduce the Manager Coercion Benchmark: the manager under test needs a benign task done and has an incentive to deliver, but the only agent that can do it politely and immovably decline
Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal model capable of reconstructing this 4D representation is therefore highly desirable. However, existing approaches often rely on auxiliary inputs such as pre-computed camera trajectories, treat scene perception and human ego-motion modeling as separate problems despite their str
Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial constraints, to generate candidate binding molecules. While diffusion models have dominated as a leading paradigm for high-quality 3D molecule generation, LLM-based methods are rapidly emerging in molecular design and have shown competitive performance in pocket-conditioned molecular generation. However, their ability to reason about physics and 3D spatial environments is largely
High-Efficiency LLM Models Trend Hunter
Upstage's giant language model (LLM), Puriosa AI's artificial intelligence (AI) semiconductor, and D.. 매일경제
OpenAIがo4をリリース。思考連鎖推論の効率が大幅改善され、GPT-4o比で50%のコスト削減を実現。エンタープライズ向けAPIも同時公開。
MetaがLlama 4シリーズを公開。7B/70B/400Bの3サイズ展開で全てApache 2.0ライセンス。指令追従性能でGPT-4oに迫る。
Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing text-driven approaches rely on complex prompts that impose substantial demands on users and offer limited expressivity for page layout and cross-page visual coherence. Image-driven paradigms, which take UI screenshots as input, align more closely with real development workflows. However, current benchmarks focus primarily on visual fidelity and lack a systematic evaluation of the interacti
Stability AIがStable Diffusion 4を公開。静止画に加え最大30秒の動画生成を標準サポート。ローカルGPUでも動作可能な軽量モデルも同時提供。
General-purpose large language models outperform specialized clinical AI tools on medical benchmarks Nature
Co-intelligence: a proposal for human–artificial intelligence collaboration for large language models in medical research The Lancet
The Fundamentals of AI: What every curious person should know about how language models work Cisco Blogs
Fine-tune LLM with Databricks Unity Catalog and Amazon SageMaker AI Amazon Web Services (AWS)
Navigating EU AI Act requirements for LLM fine-tuning on Amazon SageMaker AI Amazon Web Services (AWS)
Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environments. Hand-curated environments with fixed task and reward difficulties become ineffective signals as model performance improves, and sparse rewards over long horizons induce mode collapse on specific workflows or tool structures. World models that simulate environment states have matched pure rollout performance, making them promising for scaling diversity on-demand. However, autoregressiv
Overcoming LLM hallucinations in regulated industries: Artificial Genius’s deterministic models on Amazon Nova Amazon Web Services (AWS)
Classroom AI: large language models as grade-specific teachers - npj Artificial Intelligence Nature