LLM & AI infrastructure tools for developers
The best LLM & AI infrastructure tools built for developers, launched by indie makers and ranked by upvotes. 16 products.
- 1
Mini-AGIDynamic continual learning model trained on 8GB VRAM
mini-AGI is a byte-level language model that trains from scratch on a single 8 GB VRAM GPU and keeps learning from a continuous stream of text. Weights live on disk and are paged onto the GPU as needed, so model size is bounded by disk spac
💬 0Y▲ 277 on HNLLMs & InfrastructureResearch & DataOpen Sourceby AI Launch Team
- 2Cactus Needle 3
8-29MB automation models can match DeepSeek V4 Flash
Needle 3 is an 8-29 MB foundation model from Cactus for phones, wearables, robots, smart home devices, cars and microcontrollers. It trades general chat ability for strong tool calling, structured extraction and text embeddings, running ful
💬 0Y▲ 236 on HNLLMs & InfrastructureFreemiumby AI Launch Team
- 3LLM Attention Visualization
A visualization of the attention mechanism in LLMs.
An interactive visualization of the attention mechanism in a small large language model. Hover or tap any generated token to see which earlier tokens influenced it, with opacity scaled by attention weight across all heads and layers. Exampl
💬 0Y▲ 177 on HNResearch & DataLLMs & InfrastructureFreeby AI Launch Team
- 4JevBench
A reproducible benchmark for typed decision models
JevBench, from Benchmark Heaven, compares Jev-class decision models across intelligence, calibration, speed and cost. It sits alongside Benchmark Heaven's broader model rankings, which can be filtered by provider, lab, hosting region and da
💬 0Y▲ 153 on HNResearch & DataLLMs & InfrastructureFreeby AI Launch Team
- 5
CUA-S1A System One Model for Computer Use
Cua gives AI agents computers they can use. It provides open-source desktop automation drivers, isolated cloud desktops (Cua Fleets), local macOS VMs (Lume), CUA-S1 specialist models for computer-use decisions, and Cua Bench for evaluating
💬 0Y▲ 95 on HNAI AgentsLLMs & InfrastructureFreemiumby AI Launch Team
- 6
Lossless-memoryA personal AI memory that never summarizes
lossless-memory is a long-term memory system for a personal AI assistant that never summarizes. It keeps every raw line, timestamps everything, and searches by time first and words second, returning results in chronological order. A small i
💬 0Y▲ 67 on HNLLMs & InfrastructureProductivityOpen Sourceby AI Launch Team
- 7Jevstiller
Distill Jev into a local model – with a disagreement bound
Jevstiller distills a hosted Jev classifier into a local model, with a measured bound on how often the local model disagrees with Jev. It runs as a drop-in proxy, so batch jobs avoid a network call per answer while keeping answers close to
💬 0Y▲ 63 on HNLLMs & InfrastructureOpen Sourceby AI Launch Team
- 8Sunk Cost
How long until a local LLM rig pays for itself?
Sunk Cost estimates how long a machine for running local AI models takes to pay for itself compared with paying for an API. Pick a Mac mini, Mac Studio, DGX Spark or Strix Halo box, see which open models fit and how fast they run, and adjus
💬 0Y▲ 47 on HNLLMs & InfrastructureResearch & DataFreeby AI Launch Team
- 9
jevalsReplacing LLM judges with typed Jev decisions
jevals runs evals and guardrails for AI agents using Jev-style decision models instead of an LLM judge. All the checks for a trace, such as tool choice, groundedness, answer relevancy, prompt injection and PHI, go out as one request that co
💬 0Y▲ 47 on HNLLMs & InfrastructureAI AgentsOpen Sourceby AI Launch Team
- 10AI·rete·RAG
A Rete rule engine decides – RAG explains why
ai·rete·rag pairs a Rete rule engine with retrieval-augmented generation. The rules decide what happens, in an auditable and repeatable way, and retrieval explains why in plain language, grounded in your own documents. First shared by its
💬 0Y▲ 44 on HNLLMs & InfrastructurePaidby AI Launch Team
- 11OpenLake
Storage engine for KV-cache offload and LLM training
OpenLake is a high-performance storage engine for LLM inference and GPU training. Its Infinity Core I/O Engine delivered 6.72 GiB/s writes and 11.55 GiB/s reads in the MLPerf Storage v3.0 Llama 3.1 8B checkpointing benchmark through its S3
💬 0Y▲ 35 on HNLLMs & InfrastructureOpen Sourceby AI Launch Team
- 12
AURAA Rust agent that investigates and fixes production incidents
AURA is a production-tested SRE agent platform from Mezmo that investigates and helps fix production incidents. Specialist agents correlate traces, logs, metrics and deployment history to find root causes and recommend fixes, using the mode
💬 0Y▲ 28 on HNAI AgentsLLMs & InfrastructureOpen Sourceby AI Launch Team
- 13Lanes Link
A personal context MCP: one endpoint for every agent you use
Lanes Link is a private, self-hostable MCP server that connects your accounts, memory, skills and secrets to every AI agent you use, behind Access Profiles you manage. Connect services like Gmail, Calendar, Drive, Notion, Linear, Slack and
💬 0Y▲ 28 on HNLLMs & InfrastructureProductivityFreemiumby AI Launch Team
- 14
InstinctFlashHigh-Performance Serving Runtime for Robotics Models
InstinctFlash is a high-performance serving framework for robotics models from General Instinct (YC P26). It serves eight robotics model families, including pi05, LingBot-VLA, LingBot-VA and GR00T N1.7, through one runtime with Python and W
💬 0Y▲ 27 on HNLLMs & InfrastructureOpen Sourceby AI Launch Team
- 15OpenAPPA
Open-source deterministic guardrails that don't break agents
OpenAPPA is an information-flow policy engine that acts as a deterministic guardrail for LLM agents. It tracks how data flows instead of matching patterns, which the project says makes it fully resistant to data exfiltration from prompt inj
💬 0Y▲ 23 on HNAI AgentsLLMs & InfrastructureOpen Sourceby AI Launch Team
- 16Otis
A minimal AI agent that runs local models out of the box
Otis is a minimal AI agent from Triangl Labs that runs local models out of the box. First shared by its maker on Show HN: https://news.ycombinator.com/item?id=49696084
💬 0Y▲ 19 on HNAI AgentsLLMs & InfrastructureFreeby AI Launch Team
Related
Building one of these?
Launch it free