Search “inference”

40 matching tools

inference
Release.aiAI Models & TrainingRelease.ai deploys LLM, computer‑vision, and multimodal models with sub‑100 ms latency. It auto‑scales from zero to thousands of concurrent requests, provides enterprise‑grade security (SOC 2 Type II, private networking, end‑to‑end encryption), and offers SDKs, APIs, and real‑time monitoring.release.aiFree5.0actcast.ioAI Models & TrainingActcast is an IoT platform that runs deep‑learning inference on edge devices, detecting objects such as cats and faces locally. It reduces data transfer costs, protects privacy, and provides webhook APIs for real‑time alerts and cloud integration.actcast.ioFreeCamocopyAI Models & TrainingCamoCopy is a privacy‑first generative AI assistant offering encrypted chat, web search, and image recognition. It runs local inference on models like LLAMA, Mistral, Claude, or GPT, stores data on EU servers, auto‑deletes uploads, and supports secure workspaces with GDPR‑compliant enterprise collaboration.camocopy.comcartesia.aiAI Models & TrainingCartesia.ai is a multimodal intelligence platform that enables real-time, on-device inference with a focus on privacy and dynamic learning. It features a generative voice API for ultra-realistic audio outputs, making it suitable for diverse applications across various devices.cartesia.aiCerebrasDeveloper Tools & APIsCerebras provides a wafer-scale AI accelerator and software stack that enables single-node training of very large LLMs, high-throughput low-latency inference (GLM-4.6 at 1,000 TPS), PyTorch SDK, deployment options, and MLOps tooling.cerebras.aiFree3.9ChartAI Models & TrainingEfficient machine learning inference on-demand with lightning speed in the cloud.getcharteditor.comcirrascale.comDeveloper Tools & APIsCirrascale offers a private AI cloud that supports training and inference on AMD, Cerebras, NVIDIA, and Qualcomm accelerators. It provides zero DevOps, no data‑transfer fees, high‑bandwidth networking, and configurable multi‑GPU servers, streamlining workflows and accelerating deployment.cirrascale.comFreecortexlabs.aiAI Models & TrainingCortex is a blockchain platform that integrates AI into decentralized applications, enabling on-chain AI inference with GPU resources. It features smart contracts with machine learning, supports Solidity, and offers a collaborative ecosystem for AI model sharing.cortexlabs.aiFreedeepsense.aiAI Models & TrainingDeepSense.ai provides end‑to‑end AI solutions for enterprises, integrating large language models, retrieval‑augmented generation, MLOps, advanced computer‑vision, edge inference, and predictive analytics to deliver scalable, real‑time AI agents, co‑pilots, and maintenance optimization.deepsense.ai5.0EnergeticAIDeveloper Tools & APIsEnergeticAI is an open‑source TensorFlow.js library for Node.js, offering fast pre‑trained embeddings, text classifiers, and semantic search. It delivers sub‑4‑second cold starts and 67× faster inference in serverless functions for developers and performance.energeticai.orgfal.aiAI Models & Trainingfal.ai offers a unified API for generating images, videos, audio, and 3D models from a library of over 1,000 production‑ready assets. It provides serverless GPU inference, private deployment options, NVIDIA‑cluster fine‑tuning, SOC 2 compliance, and enterprise‑grade support.fal.ai3.7fireworks.aiAI Models & TrainingFireworks AI is a cloud‑hosted inference platform supporting code, conversational, agentic, and search workflows across text, vision, audio, and image modalities. It delivers scalable, low‑latency inference with secure RAG and serverless GPU options.fireworks.ai5.0Flyte v1.3.0Developer Tools & APIsFlyte is an open‑source Python‑based workflow platform for AI, ML, and data teams, providing self‑healing pipelines with dynamic retries, state recovery, Kubernetes autoscaling, and native integration with Spark, Ray, BigQuery, and more, enabling efficient inference and training.flyte.orgFreeGabberDeveloper Tools & APIsGabber is a cloud-based platform for developing real-time AI applications with voice, vision, and speech capabilities using a drag-and-drop interface. It supports low-latency inference, customizable inputs, and collaborative features for diverse multi-modal projects.gabber.devgpt-oss playgroundDeveloper Tools & APIsgpt-oss playground provides open-weight demos of gpt-oss-120b and 20b for infrastructure testing, distributed and on-device inference, benchmarking, API integration, and reproducible research, with adjustable reasoning levels and visible-reasoning for diagnostics. Demo-only; validate outputs.gpt-oss.comFree5.0GPUX.AIAI Models & TrainingGPUX is a serverless inference platform that delivers 1‑second cold starts and GPU‑accelerated execution for models like Stable Diffusion XL, ESRGAN, and Whisper. It supports P2P and read‑write volume access for rapid, scalable deployment on NVIDIA RTX 4090 GPUs.gpux.aiFreeGroqAI Models & TrainingGroq is an inference platform that uses custom LPU silicon for low‑latency, high‑throughput AI workloads. It supports large language and multimodal models via an OpenAI‑compatible API, with modular deployment and predictable performance for NLP, vision, and recommendation tasks.groq.com4.1hiveblockchain.comAI Models & TrainingHive Digital Technologies Ltd provides scalable cloud computing resources using industrial GPUs for AI training and inference. Its renewable energy-powered data centers enhance operational efficiency while supporting Bitcoin mining and contributing to the decentralized digital economy.hiveblockchain.comFreeHyperMinkDeveloper Tools & APIsHyperMink AI is an open‑source, privacy‑centric platform offering a modular Node.js inference server, Inferenceable, powered by llama.cpp/llamafile. It supports local model deployment, plug‑in extensions, and community contributions via GitHub for developers.hypermink.comFreeLightning AIAI Models & TrainingLightning AI is a PyTorch Lightning‑based cloud platform for training, deploying, and serving models at scale. It offers GPU workspaces, managed clusters, fractional pay‑as‑you‑go GPU capacity, inference APIs, serverless deployment, security, and integration with LitServe, LitGPT, and LLMs.lightning.ailocal.aiAI Models & Traininglocal.ai runs language models locally without GPUs. Its Rust backend keeps the binary under 10 MB and performs CPU inference with GGML quantization. A single‑click interface streams responses to a UI, while a model manager tracks, verifies, and resumes downloads.localai.appFreeMindplixAI Models & TrainingMindplix AI Labs offers a VPS platform that lets users deploy and run open‑source AI models—such as DeepSeek, ChatGPT, LTX‑2, Open‑Sora‑v2, Nvidia Nemotron—without MLOps skills. It provides unlimited inference, automatic scaling, a clean dashboard/API, and keeps all data on private servers.mindplix.comNexa.aiDeveloper Tools & APIsNexa AI offers an on‑device platform that lets developers deploy vision, audio, and text models to NPUs, GPUs, and CPUs with one line of code. The SDK supports day‑zero deployment, multimodal inference, and optimizations for mobile, automotive, and IoT devices.nexa4ai.comFreeocto.aiAI Audio & VoiceOctoAI is a generative AI framework that enhances model inference for developers and enterprises. It supports customizable models, integrates with cloud services, and offers low-latency performance for applications like image generation and voice dubbing.octo.aiFree
Snapshot mode · This page is served from a database snapshot exported on 2026-09-16, not live data. This notice disappears once the live API is connected.