Description:
LongCat-Video is a 13.6B open-source video generation model enabling Text-to-Video, Image-to-Video, and Video Continuation with efficient, long-video generation at 720p/30fps using a unified architecture and multi-reward RLHF.
Keep Calm and Read the Friendly Manual :-)
Description:
LongCat-Video is a 13.6B open-source video generation model enabling Text-to-Video, Image-to-Video, and Video Continuation with efficient, long-video generation at 720p/30fps using a unified architecture and multi-reward RLHF.
Description:
An open-source framework for audio-driven human video generation. LongCat-Video-Avatar 1.5 improves lip-sync with Whisper-Large, enhances production stability for long videos, generalizes to stylized domains, and enables faster 8-step inference.
Description:
Next.js-powered SaaS that turns text queries into AI-generated videos with a modern UI; supports cloud-based services and API integrations.
Description:
OpenMausBot is an open-source, local-first chat app where each bot is a real AI agent with its own memory, computer, and tools. It runs on your machine with optional cloud-enabled connectors (Composio, Box) and a built-in harness that coordinates multiple bots in a Telegram-style interface.
Description:
ModLens is a vision plugin for DeepSeek Harness that lets text-only AI models read pasted images, returning structured JSON evidence (OCR, layout, semantics) and enabling vision-enabled outputs across multiple hosts.
Description:
Cua is an open-source platform for agent-ready sandboxes across macOS, Windows, Linux, and containers. It provides a unified driver, sandbox API, and benchmarks for building and evaluating autonomous agents.
License: Dual license: Open Source License (AGPL-3.0) and Enterprise License (Commercial)
Description:
Kodus AI is a monorepo for AI-powered code review, enabling users to run reviews with BYOK model providers, choose from multiple LLMs, apply Kody Rules, and monitor outcomes with Cockpit. Self-host or use Kodus Cloud.
Description:
OpenDroid is an autonomous AI agent for Android that plans, executes, and verifies tasks on-device, enabling hands-free automation with multi-LLM support and on-device model management.
Description:
FaceSwap is a Python tool that uses deep learning to recognize and swap faces in photos and videos, featuring multiple models and GUI components.
Description:
Whisper is a general-purpose speech recognition model by OpenAI, capable of multilingual transcription, translation, and language identification. Transformer-based, trained on diverse audio data.