HumaniBench: A Human-Centric Benchmark for Large Multimodal Models Evaluation
Source-code: https://github.com/VectorInstitute/humaniBench
HumRights-Bench — the first benchmark for human rights reasoning in AI
JonathanHallstrom/pawnocchio: chess engine, goal is to make it strong. currently plays good chess
codedeliveryservice/Reckless: Competitive chess engine written in Rust
JevBench — Decision models for live agents
Source-code: https://jevbench.dev/#open-source
Jev Trade | Live Jev trading bot on crypto and other assets
Source-code: https://github.com/aowang-ai/jev-trade
jarrodwatts/jev-trader: One AI trade decision every Monad block. Jev on Kuru MON-USDC.
MiniMind - Train LLMs from Scratch
Source-code: https://github.com/jingyaogong/minimind
FareedKhan-dev/kimi-k3-in-c: A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
Odysseys — a benchmark for long-horizon web agents
Together Chat
World of AI Bench | AI Coding Benchmark and Model Rankings
Is Better AI - Compare AI Models Side by Side
Source-code: https://github.com/midudev/isbetter.ai
harveyai/harvey-labs: A benchmark built to evaluate and improve agent capabilities for supporting legal work
LLM Pricing — Compare 2100+ Models & 183+ Providers · 8/2026
google-deepmind/weathernext
ChatHub - GPT-5, Claude 4.5, Gemini 3 side by side
ChatAll
ai-shifu/ChatALL: Concurrently chat with ChatGPT, Bing Chat, Bard, Alpaca, Vicuna, Claude, ChatGLM, MOSS, 讯飞星火, 文心一言 and more, discover the best answers
ChatALL
Poe - Fast, Helpful AI Chat
ARC Prize - What is ARC-AGI?
ximinng/LLM4SVG: [CVPR 2025] Official implementation for "Empowering LLMs to Understand and Generate Complex Vector Graphics" https://arxiv.org/abs/2412.11102
LLM API Pricing Comparison 2026 — Cost Per Token for GPT, Claude, Gemini & More
NVIDIA Nemotron - Build Agentic AI with Multimodal Foundation Models
Humanity's Last Exam
Benchlm.ai - LLM Leaderboard 2026