While Multimodal Large Language Models (MLLMs) excel in general tasks, rigorous scientific reasoning remains challenging due to the limitations of monolithic, l
Opportunities
161–180 of 533Always filtered to Maschmeyer Group . Fit opens the breakdown — clear screening badges via Clear triage on the company page.
Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
Show HN: BillAI Bass, an AI-Powered Big Mouth Billy Bass Using Strands Agents
Show HN: Sx 2.0 – Share AI skills with your team through a Dropbox folder
The growing adoption of local inference frameworks such as Ollama has made it increasingly common for developers to run large code models on laptops and other r
Bug reports serve as task specifications for repository-level automated program repair (APR) agents, but they often describe only the observed failure and omit
Large language models (LLMs) have demonstrated strong reasoning performance, but their tendency to hallucinate limits their reliability in knowledge-intensive t
LLM agents are increasingly evaluated on multi-week decision tasks in which the state that drives cost is never directly observed. On such tasks the final cost
Large language model (LLM) coding agents are increasingly deployed to autonomously perform software engineering tasks in terminal-based environments, making the
While modern question answering (QA) systems excel on clean, schema-aligned corpora, real-world knowledge is rarely so neatly packaged. Answering questions over
Software evolves continuously, yet ensuring that a patch preserves intended behavior without re-verifying an entire codebase remains difficult. Regression verif
Recent LLM-based multi-agent urban simulators can generate semantically rich city routines, but they remain costly to scale and are often weakly validated again
Firmware rehosting executes firmware images in emulated environments such as QEMU to enable scalable dynamic analysis of Internet of Things (IoT) devices. In pr
LLM-based agents are increasingly deployed in multi-agent environments whose incentives can shape their behavior. We introduce The Energy Society, a minimal sur
Speculative decoding accelerates large language model (LLM) inference without compromising output quality. Recent parallel drafting methods further improve sing
Coding-agent benchmarks have largely measured whether agents can produce functionally correct patches, but production software also demands measurable speedups
Large language models (LLMs) have opened new opportunities for unit test generation, but executable tests do not necessarily reveal real defects. This paper stu
Structured pruning is a hardware-friendly way to compress LLMs, but it is mostly validated on multiple-choice recognition tasks, while the same compressed check
Show HN: Self-hosted voice AI agent for Asterisk/FreePBX
AI meeting assistant on Web, Desktop & Mobile (Bot optional)