Aludel is a powerful evaluation workbench designed for Phoenix apps, enabling real-time comparisons of language models like OpenAI, Anthropic, and Ollama. Track output quality, latency, token usage, and cost while managing prompts with version control. Visualize performance trends and ensure consistent evaluation across providers.

Aludel is a powerful LLM Evaluation Workbench designed specifically for Phoenix applications. This tool streamlines the process of evaluating large language models (LLMs) from various providers such as OpenAI, Anthropic, and Ollama, allowing users to compare their performance in real-time.
{{variable}} interpolation. Each edit creates an immutable new version of a prompt, ensuring clarity and consistency in prompt evolution. Tags and descriptions help in organizing prompts effectively.contains, regex, exact_match, and json_field. This feature enables tracking of pass rates and regression detection over time.To leverage Aludel effectively:
{{variable}} syntax:
Explain {{topic}} in exactly 3 sentences.
Aludel can be seamlessly integrated into any Phoenix LiveView application as a self-contained dashboard. Setting it up involves:
mix.exsconfig/config.exsFor standalone usage, Aludel's application can be set up from the standalone/ directory, allowing operation without embedding it within a Phoenix app.
Engage with the Aludel community through Discussions for queries and idea sharing, and report issues via the Issues section on GitHub.
Contributions are welcome. Please refer to the CONTRIBUTING.md for detailed instructions on how to participate in the Aludel project.
No comments yet.
Sign in to be the first to comment.