AI agents
Tools, memory, evaluation and orchestration for agents that do real work.
104 links, newest first.
- AI agentsPost on X
TheAgentCompany Benchmarks Agents on Real-World Tasks
TheAgentCompany is a benchmark for evaluating AI agents on consequential tasks spanning software development, project management, administration, and data science.
It offers engineers a benchmark for assessing agent performance across varied work tasks.
- AI agentsPost on X
ADAS Uses Search to Design and Evaluate Agentic Systems
Automated Design of Agentic Systems (ADAS) uses Meta Agent Search to propose agents from a domain description, framework code, output instructions and examples, and an archive of discovered agents. New agents are evaluated and added to the archive; the framework was tested on several reasoning…
The iterative archive-and-evaluation process is relevant to engineers exploring automated agent design.
- AI agentsPost on X
AutoToS Refines Search Components With Tests and BFS
AutoToS iteratively refines generated successor and goal-test functions using unit-test feedback. It also checks them in a breadth-first search on example problems.
The process shows how tests and search-based checks can guide automated refinement of generated components.
- AI agentsRepository
Agentic components for Llama Stack APIs
The GitHub repository contains agentic components of the Llama Stack APIs. The post highlights using Tools and Functions with Llama 3.1 models.
A relevant starting point for engineers exploring tool and function use in Llama Stack.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor