OpenClaw: Building Local Memory on DGX Spark
Watch on YouTube →
Overview
Ray Fernando details his project to build a local contextual memory system for OpenClaw using a DGX Spark, aiming to replace cloud services by indexing data from Notion, Slack, and Google Drive. He plans to use Whisper for voice-to-text, NVEmbed v2 for semantic vectors stored in Qdrant, and GLINER for entity extraction in Falcor, with QEN3 30B as the local LLM.
Key takeaways
- Ray Fernando is building a local contextual memory system for OpenClaw using a DGX Spark to replace expensive cloud services.
- The system architecture involves multiple layers: data ingestion, semantic vector creation with NVEmbed v2, entity extraction with GLINER, and reasoning with the QEN3 30B LLM.
- Ray emphasizes the importance of data ownership and consolidating data from various sources into a single, locally-processed container.
- The project is inspired by Pulse HQ's work on real-time call transcription and aims to create a context graph for agent assistance.
- Ray is using Cursor with SSH to manage the DGX Spark and automate tasks with AI agents.
Chapters
0:00
Project Overview: Building Local Contextual Memory with OpenClaw
- Ray Fernando aims to index and search all conversations and transcriptions locally using OpenClaw.
- The goal is to create a contextual memory layer that understands and comprehends data from various sources.
- Ray plans to use a DGX Spark provided by NVIDIA to run local models for data processing.
2:05
Data Sources and System Architecture
- Data sources include Notion Docs, Slack messages, meeting recordings, Google Drive, and Markdown Notes.
- The system architecture involves a contextual memory layer stored in a database for data comprehension.
- Ray's OpenClaw runs on a Mac Mini and will query the contextual memory layer for answers.
3:23
The Shift to Local Models and Data Ownership
- Local models are now capable enough to handle these workflows on-device.
- Ray has exported ChatGPT and Anthropic chats to integrate into the local system.
- Owning the data allows for reasoning and smoother business operations.
5:12
Technical Breakdown: Layers and Models
- Layer 1 is where the data lives, Layer 2 uses Whisper for voice-to-text and NVEmbed v2 for semantic vectors.
- Qdrant database stores the vectors for semantic search.
- GLINER finds people, projects, decisions, dates, and action items, stored in the Falcor database.
6:54
LLM Integration and Hardware
- QEN3 30B LLM will run on the DGX Spark to read and retrieve context for grounded answers.
- The DGX Spark has 128GB of shared local memory.
- Ray aims to unlock solutions for users spending $500-$1000/month on cloud services.
8:27
Inspiration from Pulse HQ and Real-Time Data Processing
- Inspired by a Pulse HQ talk on YouTube, Ray aims to ingest and process real-time call transcriptions.
- The system will perform document parsing, chunking, entity extraction, and relationship mapping.
- The goal is to create a context graph for agent assistance in case report triage and recommendations.
10:38
DGX Spark vs. M3 Ultra and Wait Times
- Ray addresses a comment about comparing the DGX Spark with the Mac Studio M3 Ultra.
- The wait time for a maxed-out M3 Ultra is 11-12 weeks, making it impractical for immediate comparison.
- Ray is currently focusing on the Spark due to its immediate availability.
12:31
LLM Calculator and Community Engagement
- Ray mentions an LLM calculator resource he created, but the site is currently down.
- He plans to republish the calculator, which is open source on GitHub.
- Ray encourages viewers to repost the live stream on X (Twitter).
14:59
Data Ownership and System Limitations
- Ray emphasizes the importance of owning data and exporting it from systems like Notion, ChatGPT, and Google Search.
- Current tools have limitations as they cannot communicate with each other.
- Ray aims to consolidate all data into one container for semantic search, knowledge graphs, and local model execution.
17:30
Setup Plan and Docker Containers
- Ray plans to set up Docker containers for databases and UV environments for service exposure.
- Scripts will be created to hit endpoints for embeddings and other processes.
- OpenClaw will use skills files to hit local network endpoints.
19:03
Network Topology with Tailscale VPN
- Tailscale VPN will keep everything in a secure container.
- The Mac Mini running OpenClaw will access the private network through an encrypted tunnel.
- The DGX Spark will host Docker containers, virtual environments, and the existing setup.
20:54
Safari vs. Chrome: Privacy and Security
- Ray defends Safari as the best browser for power saving, memory optimization, and security on Mac.
- He criticizes Chrome for data leakage, fingerprinting, and running shadow services.
- Safari's private browsing mode is more secure and private than Chrome's.
23:54
SSH Setup and CUDA Verification
- Ray uses Cursor to SSH into the DGX Spark and runs commands to check CUDA versions.
- He verifies CUDA 13 and NVIDIA SMI.
- Ray also checks the ARCH64 architecture and available memory.
26:39
Disabling Swap and Docker Checks
- Ray disables swap memory on the DGX Spark.
- He verifies that Docker is installed.
- Ray checks the NVIDIA Container Toolkit and Python version.
29:11
Project Recap and Data Ingestion
- Ray recaps the project goal: creating contextual memory for Clawbot with data from various sources.
- He plans to include Notion, Slack messages, Google Drive, and real-time meeting transcriptions.
- The DGX Spark will process data and create relationships in real-time.
31:31
Data Processing Layers and Model Selection
- Layer 2 will use GLNIR2 to find people, projects, and decisions, storing them in the FALCOR database.
- Whisper transcriptions will convert text to meanings for vectors, stored in Kudrant.
- Quen3 will be used locally on the Spark for retrieval.
33:47
Research and Implementation Plan
- Ray used Exa MCP and Claude for research to decide on embedding models for the DGX Spark.
- He aims to understand how these systems work and build a similar system for learning.
- The implementation plan involves setting up databases and following a research-backed approach.
38:53
Phase One: Database Setup and Swap Configuration
- Ray starts with phase one, focusing on setting up the databases.
- He turns off swap memory for unified memory safety.
- Ray uses passwords stored in Keychain to configure the system.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Ray Fernando.