> ## Documentation Index
> Fetch the complete documentation index at: https://firecrawl-docs-firecrawl-replit-connector.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# SourceSync.ai

> Firecrawl はウェブスクレイピング機能のために SourceSync.ai と連携します。

[SourceSync.ai](https://sourcesync.ai) は、独自データを用いて AI アプリケーションを構築できる Retrieval Augmented Generation as a Service のプラットフォームです。本ガイドでは、ウェブスクレイピング機能において Firecrawl を SourceSync.ai と組み合わせて利用する方法を説明します。

<div id="setup">
  ## セットアップ
</div>

1. まず、[Firecrawl ダッシュボード](https://www.firecrawl.dev/app)から Firecrawl の API キーを取得します。

2. SourceSync.ai のネームスペースで、Web スクレイピングのプロバイダーとして Firecrawl を使用するように設定します:

<CodeGroup>
  ```bash cURL theme={null}
  curl -X PATCH https://api.sourcesync.ai/v1/namespaces/YOUR_NAMESPACE_ID \
    -H "Authorization: Bearer YOUR_SOURCE_SYNC_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "webScraperConfig": {
        "provider": "FIRECRAWL",
        "apiKey": "YOUR_FIRECRAWL_API_KEY"
      }
    }'
  ```

  ```javascript JavaScript theme={null}
  const response = await fetch(
    'https://api.sourcesync.ai/v1/namespaces/YOUR_NAMESPACE_ID',
    {
      method: 'PATCH',
      headers: {
        Authorization: `Bearer ${SOURCE_SYNC_API_KEY}`,
        'Content-Type': 'application/json',
      },
      body: JSON.stringify({
        webScraperConfig: {
          provider: 'FIRECRAWL',
          apiKey: 'YOUR_FIRECRAWL_API_KEY',
        },
      }),
    }
  )
  ```

  ```python Python theme={null}
  import requests

  response = requests.patch(
      'https://api.sourcesync.ai/v1/namespaces/YOUR_NAMESPACE_ID',
      headers={
          'Authorization': f'Bearer {SOURCE_SYNC_API_KEY}',
          'Content-Type': 'application/json',
      },
      json={
          'webScraperConfig': {
              'provider': 'FIRECRAWL',
              'apiKey': 'YOUR_FIRECRAWL_API_KEY',
          },
      },
  )
  ```
</CodeGroup>

<div id="usage">
  ## 使い方
</div>

設定が完了したら、SourceSync.ai のウェブスクレイピング用エンドポイントを Firecrawl の機能と組み合わせて利用できます。主な取り込み方法は次のとおりです：

<div id="url-list-ingestion">
  ### URLリストの取り込み
</div>

特定のURLをスクレイピングする:

<CodeGroup>
  ```bash cURL theme={null}
  curl -X POST https://api.sourcesync.ai/v1/ingest/urls \
    -H "Authorization: Bearer YOUR_SOURCE_SYNC_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "namespaceId": "YOUR_NAMESPACE_ID",
      "ingestConfig": {
        "source": "URLS_LIST",
        "config": {
          "urls": [
            "https://example.com/page1",
            "https://example.com/page2"
          ],
          "scrapeOptions": {
            "includeSelectors": ["article", "main"],
            "excludeSelectors": [".navigation", ".footer"]
          }
        }
      }
    }'
  ```

  ```javascript JavaScript theme={null}
  const response = await fetch('https://api.sourcesync.ai/v1/ingest/urls', {
    method: 'POST',
    headers: {
      Authorization: `Bearer ${SOURCE_SYNC_API_KEY}`,
      'Content-Type': 'application/json',
    },
    body: JSON.stringify({
      namespaceId: 'YOUR_NAMESPACE_ID',
      ingestConfig: {
        source: 'URLS_LIST',
        config: {
          urls: ['https://example.com/page1', 'https://example.com/page2'],
          scrapeOptions: {
            includeSelectors: ['article', 'main'],
            excludeSelectors: ['.navigation', '.footer'],
          },
        },
      },
    }),
  })
  ```

  ```python Python theme={null}
  response = requests.post(
      'https://api.sourcesync.ai/v1/ingest/urls',
      headers={
          'Authorization': f'Bearer {SOURCE_SYNC_API_KEY}',
          'Content-Type': 'application/json',
      },
      json={
          'namespaceId': 'YOUR_NAMESPACE_ID',
          'ingestConfig': {
              'source': 'URLS_LIST',
              'config': {
                  'urls': ['https://example.com/page1', 'https://example.com/page2'],
                  'scrapeOptions': {
                      'includeSelectors': ['article', 'main'],
                      'excludeSelectors': ['.navigation', '.footer'],
                  },
              },
          },
      },
  )
  ```
</CodeGroup>

<div id="website-crawling">
  ### ウェブサイトのクロール
</div>

カスタムルールでウェブサイト全体をクロールする:

<CodeGroup>
  ```bash cURL theme={null}
  curl -X POST https://api.sourcesync.ai/v1/ingest/website \
    -H "Authorization: Bearer YOUR_SOURCE_SYNC_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "namespaceId": "YOUR_NAMESPACE_ID",
      "ingestConfig": {
        "source": "WEBSITE",
        "config": {
          "url": "https://example.com",
          "maxDepth": 3,
          "maxLinks": 100,
          "includePaths": ["/docs", "/blog"],
          "excludePaths": ["/admin"],
          "scrapeOptions": {
            "includeSelectors": ["article", "main"],
            "excludeSelectors": [".navigation", ".footer"]
          }
        }
      }
    }'
  ```

  ```javascript JavaScript theme={null}
  const response = await fetch('https://api.sourcesync.ai/v1/ingest/website', {
    method: 'POST',
    headers: {
      Authorization: `Bearer ${SOURCE_SYNC_API_KEY}`,
      'Content-Type': 'application/json',
    },
    body: JSON.stringify({
      namespaceId: 'YOUR_NAMESPACE_ID',
      ingestConfig: {
        source: 'WEBSITE',
        config: {
          url: 'https://example.com',
          maxDepth: 3,
          maxLinks: 100,
          includePaths: ['/docs', '/blog'],
          excludePaths: ['/admin'],
          scrapeOptions: {
            includeSelectors: ['article', 'main'],
            excludeSelectors: ['.navigation', '.footer'],
          },
        },
      },
    }),
  })
  ```

  ```python Python theme={null}
  response = requests.post(
      'https://api.sourcesync.ai/v1/ingest/website',
      headers={
          'Authorization': f'Bearer {SOURCE_SYNC_API_KEY}',
          'Content-Type': 'application/json',
      },
      json={
          'namespaceId': 'YOUR_NAMESPACE_ID',
          'ingestConfig': {
              'source': 'WEBSITE',
              'config': {
                  'url': 'https://example.com',
                  'maxDepth': 3,
                  'maxLinks': 100,
                  'includePaths': ['/docs', '/blog'],
                  'excludePaths': ['/admin'],
                  'scrapeOptions': {
                      'includeSelectors': ['article', 'main'],
                      'excludeSelectors': ['.navigation', '.footer'],
                  },
              },
          },
      },
  )
  ```
</CodeGroup>

<div id="sitemap-processing">
  ### サイトマップの処理
</div>

サイトマップ内のすべてのURLを処理します。

<CodeGroup>
  ```bash cURL theme={null}
  curl -X POST https://api.sourcesync.ai/v1/ingest/sitemap \
    -H "Authorization: Bearer YOUR_SOURCE_SYNC_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "namespaceId": "YOUR_NAMESPACE_ID",
      "ingestConfig": {
        "source": "SITEMAP",
        "config": {
          "url": "https://example.com/sitemap.xml",
          "scrapeOptions": {
            "includeSelectors": ["article", "main"],
            "excludeSelectors": [".navigation", ".footer"]
          }
        }
      }
    }'
  ```

  ```javascript JavaScript theme={null}
  const response = await fetch('https://api.sourcesync.ai/v1/ingest/sitemap', {
    method: 'POST',
    headers: {
      Authorization: `Bearer ${SOURCE_SYNC_API_KEY}`,
      'Content-Type': 'application/json',
    },
    body: JSON.stringify({
      namespaceId: 'YOUR_NAMESPACE_ID',
      ingestConfig: {
        source: 'SITEMAP',
        config: {
          url: 'https://example.com/sitemap.xml',
          scrapeOptions: {
            includeSelectors: ['article', 'main'],
            excludeSelectors: ['.navigation', '.footer'],
          },
        },
      },
    }),
  })
  ```

  ```python Python theme={null}
  response = requests.post(
      'https://api.sourcesync.ai/v1/ingest/sitemap',
      headers={
          'Authorization': f'Bearer {SOURCE_SYNC_API_KEY}',
          'Content-Type': 'application/json',
      },
      json={
          'namespaceId': 'YOUR_NAMESPACE_ID',
          'ingestConfig': {
              'source': 'SITEMAP',
              'config': {
                  'url': 'https://example.com/sitemap.xml',
                  'scrapeOptions': {
                      'includeSelectors': ['article', 'main'],
                      'excludeSelectors': ['.navigation', '.footer'],
                  },
              },
          },
      },
  )
  ```
</CodeGroup>

<div id="features">
  ## 機能
</div>

Firecrawl を SourceSync.ai と併用すると、次の機能を利用できます:

* JavaScript レンダリング対応
* 自動レート制限
* CSS セレクターによるコンテンツ抽出
* 深さを制御した再帰クロール
* サイトマップ処理

<div id="resources">
  ## リソース
</div>

* [SourceSync.ai ドキュメント](https://sourcesync.ai)
* [Webスクレイピングガイド](https://sourcesync.ai/web-scraping)
* [APIリファレンス](https://sourcesync.ai/api-reference/data-ingestion#ingest-urls)

サポートが必要な場合:

* メール: [support@sourcesync.ai](mailto:support@sourcesync.ai)
* Discord: [コミュニティに参加](https://discord.gg/Fx3GnFKnRT)
