Sign inLaunch

LLM & AI infrastructure tools for developers

The best LLM & AI infrastructure tools built for developers, launched by indie makers and ranked by upvotes. 16 products.

  1. 1
    Mini-AGI

    Dynamic continual learning model trained on 8GB VRAM

    mini-AGI is a byte-level language model that trains from scratch on a single 8 GB VRAM GPU and keeps learning from a continuous stream of text. Weights live on disk and are paged onto the GPU as needed, so model size is bounded by disk spac

    💬 0Y▲ 277 on HNLLMs & InfrastructureResearch & DataOpen Sourceby AI Launch Team

  2. 2
    Cactus Needle 3

    8-29MB automation models can match DeepSeek V4 Flash

    Needle 3 is an 8-29 MB foundation model from Cactus for phones, wearables, robots, smart home devices, cars and microcontrollers. It trades general chat ability for strong tool calling, structured extraction and text embeddings, running ful

    💬 0Y▲ 236 on HNLLMs & InfrastructureFreemiumby AI Launch Team

  3. 3
    LLM Attention Visualization

    A visualization of the attention mechanism in LLMs.

    An interactive visualization of the attention mechanism in a small large language model. Hover or tap any generated token to see which earlier tokens influenced it, with opacity scaled by attention weight across all heads and layers. Exampl

    💬 0Y▲ 177 on HNResearch & DataLLMs & InfrastructureFreeby AI Launch Team

  4. 4
    JevBench

    A reproducible benchmark for typed decision models

    JevBench, from Benchmark Heaven, compares Jev-class decision models across intelligence, calibration, speed and cost. It sits alongside Benchmark Heaven's broader model rankings, which can be filtered by provider, lab, hosting region and da

    💬 0Y▲ 153 on HNResearch & DataLLMs & InfrastructureFreeby AI Launch Team

  5. 5
    CUA-S1

    A System One Model for Computer Use

    Cua gives AI agents computers they can use. It provides open-source desktop automation drivers, isolated cloud desktops (Cua Fleets), local macOS VMs (Lume), CUA-S1 specialist models for computer-use decisions, and Cua Bench for evaluating

    💬 0Y▲ 95 on HNAI AgentsLLMs & InfrastructureFreemiumby AI Launch Team

  6. 6
    Lossless-memory

    A personal AI memory that never summarizes

    lossless-memory is a long-term memory system for a personal AI assistant that never summarizes. It keeps every raw line, timestamps everything, and searches by time first and words second, returning results in chronological order. A small i

    💬 0Y▲ 67 on HNLLMs & InfrastructureProductivityOpen Sourceby AI Launch Team

  7. 7
    Jevstiller

    Distill Jev into a local model – with a disagreement bound

    Jevstiller distills a hosted Jev classifier into a local model, with a measured bound on how often the local model disagrees with Jev. It runs as a drop-in proxy, so batch jobs avoid a network call per answer while keeping answers close to

    💬 0Y▲ 63 on HNLLMs & InfrastructureOpen Sourceby AI Launch Team

  8. 8
    Sunk Cost

    How long until a local LLM rig pays for itself?

    Sunk Cost estimates how long a machine for running local AI models takes to pay for itself compared with paying for an API. Pick a Mac mini, Mac Studio, DGX Spark or Strix Halo box, see which open models fit and how fast they run, and adjus

    💬 0Y▲ 47 on HNLLMs & InfrastructureResearch & DataFreeby AI Launch Team

  9. 9
    jevals

    Replacing LLM judges with typed Jev decisions

    jevals runs evals and guardrails for AI agents using Jev-style decision models instead of an LLM judge. All the checks for a trace, such as tool choice, groundedness, answer relevancy, prompt injection and PHI, go out as one request that co

    💬 0Y▲ 47 on HNLLMs & InfrastructureAI AgentsOpen Sourceby AI Launch Team

  10. 10
    AI·rete·RAG

    A Rete rule engine decides – RAG explains why

    ai·rete·rag pairs a Rete rule engine with retrieval-augmented generation. The rules decide what happens, in an auditable and repeatable way, and retrieval explains why in plain language, grounded in your own documents. First shared by its

    💬 0Y▲ 44 on HNLLMs & InfrastructurePaidby AI Launch Team

  11. 11
    OpenLake

    Storage engine for KV-cache offload and LLM training

    OpenLake is a high-performance storage engine for LLM inference and GPU training. Its Infinity Core I/O Engine delivered 6.72 GiB/s writes and 11.55 GiB/s reads in the MLPerf Storage v3.0 Llama 3.1 8B checkpointing benchmark through its S3

    💬 0Y▲ 35 on HNLLMs & InfrastructureOpen Sourceby AI Launch Team

  12. 12
    AURA

    A Rust agent that investigates and fixes production incidents

    AURA is a production-tested SRE agent platform from Mezmo that investigates and helps fix production incidents. Specialist agents correlate traces, logs, metrics and deployment history to find root causes and recommend fixes, using the mode

    💬 0Y▲ 28 on HNAI AgentsLLMs & InfrastructureOpen Sourceby AI Launch Team

  13. 13
    Lanes Link

    A personal context MCP: one endpoint for every agent you use

    Lanes Link is a private, self-hostable MCP server that connects your accounts, memory, skills and secrets to every AI agent you use, behind Access Profiles you manage. Connect services like Gmail, Calendar, Drive, Notion, Linear, Slack and

    💬 0Y▲ 28 on HNLLMs & InfrastructureProductivityFreemiumby AI Launch Team

  14. 14
    InstinctFlash

    High-Performance Serving Runtime for Robotics Models

    InstinctFlash is a high-performance serving framework for robotics models from General Instinct (YC P26). It serves eight robotics model families, including pi05, LingBot-VLA, LingBot-VA and GR00T N1.7, through one runtime with Python and W

    💬 0Y▲ 27 on HNLLMs & InfrastructureOpen Sourceby AI Launch Team

  15. 15
    OpenAPPA

    Open-source deterministic guardrails that don't break agents

    OpenAPPA is an information-flow policy engine that acts as a deterministic guardrail for LLM agents. It tracks how data flows instead of matching patterns, which the project says makes it fully resistant to data exfiltration from prompt inj

    💬 0Y▲ 23 on HNAI AgentsLLMs & InfrastructureOpen Sourceby AI Launch Team

  16. 16
    Otis

    A minimal AI agent that runs local models out of the box

    Otis is a minimal AI agent from Triangl Labs that runs local models out of the box. First shared by its maker on Show HN: https://news.ycombinator.com/item?id=49696084

    💬 0Y▲ 19 on HNAI AgentsLLMs & InfrastructureFreeby AI Launch Team

Related

Building one of these?

Launch it free