# Index Google Drive documents with OpenAI embeddings and Pinecone

> Quick overview This workflow keeps a Pinecone knowledge base in sync with a Google Drive folder by extracting text from supported files, generating OpenAI embeddings, and upserting vectors with metadata, while emailing an admin via Gmail when unsupported formats are found. How it works 1. Runs manually, on a daily schedule (3:00 AM), or when a new file is created in a specific Google Drive folder. 2. Lists all files in the configured Google Drive folder, skips temporary conversion/OCR files, and processes the remaining files one at a time. 3. Checks an n8n Data Table to skip files that are already indexed and unchanged since the last recorded modified time. 4. Detects each file’s MIME type and extracts content by exporting Google Docs/Sheets, extracting text from PDFs (or OCR-converting scanned PDFs), parsing Excel files, describing images with OpenAI Vision, or converting DOCX/PPTX to Google formats before export. 5. Splits the extracted content into chunks, creates OpenAI embeddings,

- **Domain**: `n8n-workflow-19424.n8n.io`
- **Provider**: n8n.io (https://n8n.io)
- **Kind**: workflow
- **Live-responds (last probe)**: None
- **Signed card**: False
- **Streaming**: False
- **Quality score**: 40%

## URLs
- Agent card: https://n8n.io/workflows/19424
- Page (HTML): https://agenstry.com/agents/n8n-workflow-19424.n8n.io
- Documentation: https://n8n.io/workflows/19424
