> ## Documentation Index
> Fetch the complete documentation index at: https://supermemory-capy-add-llmstxt-summary-and.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Supported content types

> All the content formats Supermemory can ingest and process

Supermemory automatically extracts and indexes content from various formats. There are two entry points: `client.add()` for text and URLs, `client.documents.uploadFile()` for actual files. See [Add Memories](/ingestion/add-memories) to learn how to ingest content via the API.

## Text content

Raw text, conversations, notes, or any string content.

```typescript theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
await client.add({
  content: "User prefers dark mode and uses vim keybindings",
  containerTag: "user_123"
});
```

**Best for:** Chat messages, user preferences, notes, logs, transcripts.

***

## URLs & Web pages

Send a URL and Supermemory fetches, extracts, and indexes the content.

```typescript theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
await client.add({
  content: "https://docs.example.com/api-reference",
  containerTag: "documentation"
});
```

**Extracts:** Article text, headings, metadata. Strips navigation, ads, boilerplate. URL extraction is powered by [Markdowner](https://md.dhr.wtf).

***

## Documents

### PDF

Files are binary, so they go through `uploadFile`, not `add` — pass a stream, not base64:

```typescript theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
import fs from 'fs';

await client.documents.uploadFile({
  file: fs.createReadStream('report.pdf'),
  containerTag: "user_123",
  metadata: JSON.stringify({ title: "Q4 Financial Report" })
});
```

**Extracts:** Text, tables, headers. OCR for scanned documents.

#### Advanced document extraction

**Scale and Enterprise** include advanced extraction for PDFs: page-by-page OCR with descriptions of figures and diagrams merged into the extracted text. This helps make the visual content in reports, research papers, and technical documents searchable alongside their text.

**Free, Pro, and Max** continue to support standard PDF extraction, including OCR for scanned documents. The advanced extraction tier applies to PDFs; it does not change support for other file types.

Use the same `uploadFile` call shown above. Supermemory selects the extraction tier automatically from your organization's plan; no additional request parameter is needed. If advanced extraction cannot process a PDF, Supermemory falls back to standard extraction, so figure and diagram descriptions are not guaranteed for every document.

See [Billing & usage](/overview/billing#feature-availability) for plan availability.

### Microsoft Office

Word, Excel, and PowerPoint files upload the same way — Supermemory detects the type from the file itself:

```typescript theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
await client.documents.uploadFile({
  file: fs.createReadStream('roadmap.docx'),
  containerTag: "user_123",
  metadata: JSON.stringify({ title: "Product Roadmap" })
});
```

### Google Workspace

Automatically handled via [Google Drive connector](/connectors/google-drive):

* Google Docs
* Google Sheets
* Google Slides

***

## Code & markdown

Both are plain text, so they go through `add` like any other string content — no file upload needed:

```typescript theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
// Markdown
await client.add({
  content: markdownContent,
  containerTag: "user_123",
  metadata: { title: "README.md" }
});

// Code (language auto-detected)
await client.add({
  content: codeContent,
  containerTag: "user_123",
  metadata: { language: "typescript" }
});
```

**Extracts:** Structure, headings, code blocks with syntax awareness.

Code is chunked using [code-chunk](https://github.com/supermemoryai/code-chunk), which understands AST boundaries to keep functions, classes, and logical blocks intact. See [Super RAG](/concepts/super-rag) for how Supermemory optimizes chunking for each content type.

***

## Images

`fileType: "image"` and `mimeType` are both required so Supermemory knows exactly how to process it:

```typescript theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
await client.documents.uploadFile({
  file: fs.createReadStream('diagram.png'),
  fileType: "image",
  mimeType: "image/png",
  containerTag: "user_123",
  metadata: JSON.stringify({ title: "Architecture Diagram" })
});
```

**Extracts:** OCR text, visual descriptions, diagram interpretations.

**Supported:** PNG, JPG, JPEG, WebP, GIF

***

## Audio & video

Video has a dedicated `fileType`; audio is uploaded the same way and detected from the file itself:

```typescript theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
// Video
await client.documents.uploadFile({
  file: fs.createReadStream('demo.mp4'),
  fileType: "video",
  mimeType: "video/mp4",
  containerTag: "user_123",
  metadata: JSON.stringify({ title: "Product Demo" })
});

// Audio
await client.documents.uploadFile({
  file: fs.createReadStream('call-recording.mp3'),
  mimeType: "audio/mpeg",
  containerTag: "user_123",
  metadata: JSON.stringify({ title: "Customer Call Recording" })
});
```

**Extracts:** Transcription, speaker detection, topic segmentation.

**Supported:** MP3, WAV, M4A, MP4, WebM

***

## Structured data

JSON and CSV are text — stringify and send them through `add()`, no file upload needed.

### JSON

```typescript theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
await client.add({
  content: JSON.stringify(userData),
  containerTag: "user_123",
  metadata: { title: "User Profile Data", format: "json" }
});
```

### CSV

```typescript theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
await client.add({
  content: csvContent,
  containerTag: "user_123",
  metadata: { title: "Sales Data Q4", format: "csv" }
});
```

***

## File upload

For any binary file, use `uploadFile` — it accepts a stream, not base64:

```typescript theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
import fs from 'fs';

await client.documents.uploadFile({
  file: fs.createReadStream('./document.pdf'),
  containerTag: "user_123",
  metadata: JSON.stringify({ title: "document.pdf" })
});
```

No Node `fs` access? `uploadFile` also accepts a web `File`, a `fetch` `Response`, or the SDK's `toFile` helper:

```typescript theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
import Supermemory, { toFile } from 'supermemory';

await client.documents.uploadFile({ file: new File(['my bytes'], 'file') });
await client.documents.uploadFile({ file: await fetch('https://somesite/file') });
await client.documents.uploadFile({ file: await toFile(Buffer.from('my bytes'), 'file') });
```

***

## Auto-Detection

`add()` tells URLs and plain text apart on its own — no extra flag needed:

```typescript theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
// URL detected automatically
await client.add({ content: "https://example.com/page" });

// Plain text detected automatically
await client.add({ content: "User said they prefer email contact" });
```

For files, `uploadFile` detects type from the file itself in most cases. `fileType` only exists to force specific processing — and it's required (along with `mimeType`) for images and video.

***

## Content limits

| Type | Max Size |
| - | - |
| Text | 1MB |
| Files | 50MB |
| URLs | Fetched content up to 10MB |

**Typical processing time:** text is near-instant; PDFs take 1-5s; images 2-10s; video 10s+; webpages 1-3s. Text content is chunked at the sentence level with a 2-sentence overlap between chunks.

<Tip>
  For large files, consider chunking or using [connectors](/connectors/overview) for automatic sync.
</Tip>

***

## Next steps

<CardGroup cols={2}>
  <Card title="Add memories" icon="https://mintcdn.com/supermemory-capy-add-llmstxt-summary-and/LIMkcglt81IfjBVR/icons/hugeicons/plus-sign.svg?fit=max&auto=format&n=LIMkcglt81IfjBVR&q=85&s=e61184a2a5a45db092cdb1945c0c97c0" href="/ingestion/add-memories" width="24" height="24" data-path="icons/hugeicons/plus-sign.svg">
    Upload content via the API
  </Card>

  <Card title="Super RAG" icon="https://mintcdn.com/supermemory-capy-add-llmstxt-summary-and/LIMkcglt81IfjBVR/icons/hugeicons/flash.svg?fit=max&auto=format&n=LIMkcglt81IfjBVR&q=85&s=3cb7e337be4be7f18dd98e01022d8ae4" href="/concepts/super-rag" width="24" height="24" data-path="icons/hugeicons/flash.svg">
    How content is chunked and indexed
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.