> ## Documentation Index
> Fetch the complete documentation index at: https://docs.closient.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Open a retailer catalog import job

> Open a job for one catalog source and return its id. The job starts in `staging`: stage rows with `POST .../jobs/{job_id}/rows`, then run it with `POST .../jobs/{job_id}/start`.

The run merges rows into the shared catalog GTIN-first — one product per GTIN across every retailer, one listing per retailer that carries it, and per-field provenance recording which retailer supplied what. A populated field is displaced only by a strictly higher measured confidence, which is what makes a rerun of the same feed inert and keeps the two feeds' disagreements resolved the same way whatever order they run in.

`dry_run` defaults to **true**. Send `dry_run: false` to commit.



## OpenAPI

````yaml /openapi/openapi-products.json post /products/api/v1/import/retailer-catalog/jobs
openapi: 3.1.0
info:
  title: Products API
  version: 1.0.0
  description: >
    Look up, claim, and browse GTINs in the Closient product repository.


    ## Authentication


    All endpoints require an API key passed via the `X-API-Key` HTTP header,
    unless otherwise noted.


    ```

    X-API-Key: csb_<body>_<checksum>

    ```


    Generate API keys in **Settings > API Keys** in your dashboard, or via the
    Account API.

    Session-based (cookie) authentication is also accepted for browser-based
    access.


    ## Rate Limits


    | Tier        | Requests / minute | Requests / day |

    |-------------|-------------------|----------------|

    | Default     | 300               | 10,000         |

    | Custom      | Contact us        | Contact us     |


    Rate-limit headers are included on every response so callers can
    self-throttle without

    hitting our 429s ("informed governor"):


    - `RateLimit-Policy` — every active window, e.g. `300;w=60, 10000;w=86400`

    - `RateLimit-Limit` — quota for the **most-restrictive** currently-active
    window

    - `RateLimit-Remaining` — requests left in that window

    - `RateLimit-Reset` — seconds until that window resets (relative; clock-skew
    safe)


    Legacy `X-RateLimit-*` aliases are also emitted for back-compat.
    `X-RateLimit-Reset`

    keeps the absolute Unix-timestamp shape to avoid breaking existing
    consumers.


    When rate-limited, you receive `429 Too Many Requests` with a
    `retry_after_seconds` field

    in the error envelope and a `Retry-After` header.


    ## Pagination


    List endpoints return paginated results in this envelope:


    ```json

    {
      "data": [...],
      "pagination": {
        "page": 1,
        "page_size": 25,
        "total_count": 342,
        "total_pages": 14,
        "has_next": true,
        "has_previous": false
      }
    }

    ```


    Use `?page=2&page_size=50` query parameters. Maximum page size is 100.


    ## Error Responses


    All errors conform to [RFC 9457 Problem
    Details](https://www.rfc-editor.org/rfc/rfc9457)

    with `Content-Type: application/problem+json`:


    ```json

    {
      "type": "https://closient.com/docs/errors/not_found",
      "title": "Not Found",
      "status": 404,
      "detail": "The requested resource was not found.",
      "error_code": "not_found",
      "retryable": false,
      "timestamp": "2026-03-31T12:00:00+00:00"
    }

    ```


    Common error codes: `unauthorized` (401), `forbidden` (403), `not_found`
    (404),

    `validation_error` (422), `rate_limited` (429), `internal_error` (500).
  termsOfService: https://www.closient.com/terms/
servers:
  - url: https://www.closient.com
security: []
tags:
  - name: Products
    description: Look up, claim, and browse products and trade items.
  - name: Product Group
    description: Manage GDSN packaging hierarchy relationships.
  - name: Import
    description: >-
      Bulk import products, batch/lots (GS1 AI 10) and serial numbers (GS1 AI
      21) from CSV, TSV or XLSX files, with a downloadable template per entity.
  - name: QR
    description: Generate QR codes encoding GS1 Digital Link URLs.
  - name: Codes
    description: Generate GS1 DataMatrix and 1D barcodes (EAN/UPC/ITF-14/Code128).
  - name: Digital Link
    description: >-
      Parse GS1 Digital Link URIs into structured AIs (GTIN, lot, expiry,
      serial).
  - name: Labels
    description: >-
      Bulk label-export jobs: ZIPs of QR / DataMatrix / 1D symbols and multi-up
      sheet PDFs, async via Celery with status polling.
  - name: Lots
    description: >-
      Lot generator v2: format templates with persistent counters,
      reserve/commit/discard runs, soft/hard collision handling, and an async
      Celery path with SSE progress.
  - name: Retailer Catalog Import
    description: >-
      Server-side retailer catalog import jobs: stage normalised rows, then run
      the GTIN-keyed cross-retailer merge with a dry-run mode, per-field
      provenance and a quarantine report for every row refused.
  - name: HRI Presets
    description: >-
      Saved, named per-organisation Human Readable Interpretation (HRI)
      configurations — placement, layout, notation and typography. Referenceable
      by name from the QR generation endpoint, with an org default applied when
      no options or preset are supplied.
  - name: Product Links
    description: >-
      Generic, GS1-link-type-keyed brand CTA links (e.g. reviews, promotions,
      loyalty programs, homepage) attached to a product or a brand and rendered
      on the hosted page.
externalDocs:
  description: Closient Documentation
  url: https://docs.closient.com
paths:
  /products/api/v1/import/retailer-catalog/jobs:
    post:
      tags:
        - Retailer Catalog Import
      summary: Open a retailer catalog import job
      description: >-
        Open a job for one catalog source and return its id. The job starts in
        `staging`: stage rows with `POST .../jobs/{job_id}/rows`, then run it
        with `POST .../jobs/{job_id}/start`.


        The run merges rows into the shared catalog GTIN-first — one product per
        GTIN across every retailer, one listing per retailer that carries it,
        and per-field provenance recording which retailer supplied what. A
        populated field is displaced only by a strictly higher measured
        confidence, which is what makes a rerun of the same feed inert and keeps
        the two feeds' disagreements resolved the same way whatever order they
        run in.


        `dry_run` defaults to **true**. Send `dry_run: false` to commit.
      operationId: apps_products_api_retailer_catalog_create_catalog_job
      parameters: []
      requestBody:
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/CatalogJobCreateIn'
        required: true
      responses:
        '200':
          description: OK
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/CatalogJobOut'
        '400':
          description: Bad Request
          content:
            application/problem+json:
              schema:
                $ref: '#/components/schemas/ErrorOut'
        '401':
          description: Unauthorized
          content:
            application/problem+json:
              schema:
                $ref: '#/components/schemas/ErrorOut'
        '403':
          description: Forbidden
          content:
            application/problem+json:
              schema:
                $ref: '#/components/schemas/ErrorOut'
        '404':
          description: Not Found
          content:
            application/problem+json:
              schema:
                $ref: '#/components/schemas/ErrorOut'
        '405':
          description: Method Not Allowed
          content:
            application/problem+json:
              schema:
                $ref: '#/components/schemas/ErrorOut'
        '422':
          description: Unprocessable Content
          content:
            application/problem+json:
              schema:
                $ref: '#/components/schemas/ErrorOut'
        '429':
          description: Too Many Requests
          content:
            application/problem+json:
              schema:
                $ref: '#/components/schemas/ErrorOut'
      security:
        - APIKeyHeaderAuth: []
        - OAuthTokenAuth: []
        - CookieGatedSessionAuth: []
components:
  schemas:
    CatalogJobCreateIn:
      description: Flags for one catalog import run.
      examples:
        - batch_size: 500
          dry_run: true
          fetched_at: '2025-12-01T00:00:00Z'
          reclaim_unattributed: true
          skip_ingredients: false
          source: ulta_catalog
          source_file: ulta_catalog_2025-12.psv
      properties:
        source:
          $ref: '#/components/schemas/CatalogImportSourceEnum'
          description: >-
            Catalog source this run merges under. One of: `sephora_catalog`,
            `ulta_catalog`. Not free text — the source decides the canonical
            retailer, its storefront, and the measured per-field confidence
            table that resolves every contested field, so an unregistered value
            is rejected with 422 rather than guessed at. The registered set is
            asserted against the engine's own registry in the C-5966 tests.
        source_file:
          default: ''
          description: >-
            Name of the feed drop these rows came from. Recorded as listing
            provenance, so a listing can be traced to the file it came from.
          maxLength: 255
          title: Source File
          type: string
        fetched_at:
          anyOf:
            - format: date-time
              type: string
            - type: 'null'
          description: >-
            When the feed drop was fetched — provenance about the **data**, not
            the run. Send the drop's own timestamp; omitting it stamps the run's
            start time, which misdates a feed imported weeks after it was
            published.
          title: Fetched At
        dry_run:
          default: true
          description: >-
            Run the real merge and roll back each batch. Default. The counts are
            directly comparable with a committed run's, so this is a real
            rehearsal rather than an estimate. Note that quarantine rows roll
            back too and are reported from the run's in-memory record instead.
          title: Dry Run
          type: boolean
        reclaim_unattributed:
          default: true
          description: >-
            Allow this run to claim fields on products that carry no recorded
            provenance — rows from the pre-C-4096 importers, which overwrote
            rather than enriched and left no `field_sources`. Restricted to
            products still marked imported/seeded, so a brand-claimed product or
            one sourced from a crawl is never touched. On by default (C-5960); a
            dry run reports exactly how many fields it would claim, and the flag
            can still be set to false for a specific run.
          title: Reclaim Unattributed
          type: boolean
        skip_ingredients:
          default: false
          description: Skip ingredient-list parsing, the slowest part of a run.
          title: Skip Ingredients
          type: boolean
        batch_size:
          default: 500
          description: >-
            Rows per transaction window inside the engine. Each window costs a
            fixed handful of queries regardless of size, so the default rarely
            needs changing. It is also the unit a run resumes from after an
            interruption, so a smaller window loses less work and costs more
            checkpoint writes.
          maximum: 5000
          minimum: 1
          title: Batch Size
          type: integer
        chunk_rows:
          anyOf:
            - maximum: 12000
              minimum: 1
              type: integer
            - type: 'null'
          description: >-
            Staged rows merged by one background chunk task. The run is a
            sequential chain of these. Omit it and the size is chosen from
            `dry_run`: 10,000 for a dry run, 2,000 for a commit. The two modes
            are not the same amount of work per row — a commit writes every row
            and parses every ingredient label inline, measured at 14-25 rows/s
            against a dry run's ~103 — and sizing both from the dry run is what
            killed the first full-feed commit run. Send a value only if you have
            measured your own throughput; the maximum is 12,000.
          title: Chunk Rows
      required:
        - source
      title: CatalogJobCreateIn
      type: object
    CatalogJobOut:
      description: A job's current state. The same shape from open, append, start and poll.
      examples:
        - chunk_rows: 10000
          chunks_done: 25
          chunks_total: 25
          completed_at: '2026-09-11T07:41:00Z'
          dry_run: true
          job_id: b2c3d4e5-f678-9012-abcd-ef2345678901
          poll_url: >-
            /products/api/v1/import/retailer-catalog/jobs/b2c3d4e5-f678-9012-abcd-ef2345678901
          processed_rows: 248566
          reclaim_unattributed: true
          result:
            dual_listed: 14300
            listings_created: 247566
            products_created: 233241
            products_enriched: 14300
            quarantine:
              - detail: net_content is 117 characters, limit 100
                raw_gtin: '3378872412345'
                reason: field_too_long
                row_number: 18342
                source_reference: '2598765'
            quarantine_reasons:
              field_too_long: 1
            quarantined: 1
            rows_read: 248566
          source: ulta_catalog
          source_file: ulta_catalog_2025-12.psv
          staged_rows: 248566
          started_at: '2026-09-11T07:10:00Z'
          status: completed
      properties:
        job_id:
          description: Identifier of the job. Use it on the row, start and poll calls.
          format: shortuuid
          maxLength: 22
          minLength: 22
          pattern: ^[23456789ABCDEFGHJKLMNPQRSTUVWXYZabcdefghijkmnopqrstuvwxyz]{22}$
          title: Job Id
          type: string
        status:
          description: >-
            `staging` — accepting rows; the run has not started. `queued` —
            start accepted, waiting for a worker. `running` — the merge is
            executing; `chunks_done` and `result` advance as it goes.
            `completed` — finished; `result` carries the whole feed's census.
            `failed` — the run gave up; `error` says why, `result` carries the
            census of the work that did finish, and `POST .../start` resumes
            from where it stopped.
          title: Status
          type: string
        source:
          description: Catalog source this run merges under.
          title: Source
          type: string
        source_file:
          default: ''
          description: Feed drop name recorded as listing provenance.
          title: Source File
          type: string
        dry_run:
          description: Whether this run rolls back each batch.
          title: Dry Run
          type: boolean
        reclaim_unattributed:
          description: Whether this run may claim unattributed fields.
          title: Reclaim Unattributed
          type: boolean
        staged_rows:
          description: Rows accepted into the job so far.
          minimum: 0
          title: Staged Rows
          type: integer
        chunk_rows:
          description: >-
            Staged rows merged by one chunk task. The run is a sequential chain
            of these, so no single task approaches the worker's time limit. Set
            at open time from `dry_run` unless you named one: 10,000 for a dry
            run, 2,000 for a commit, because a commit does an order of magnitude
            more work per row.
          minimum: 1
          title: Chunk Rows
          type: integer
        chunks_total:
          description: >-
            Chunks this job's staged rows divide into. Zero until rows are
            staged.
          minimum: 0
          title: Chunks Total
          type: integer
        chunks_done:
          description: >-
            Chunks fully merged and recorded. A resumed run continues from here,
            so these rows are never merged twice.
          minimum: 0
          title: Chunks Done
          type: integer
        processed_rows:
          description: >-
            Staged rows already merged. A floor, not an estimate: rows counted
            here are recorded and survive a restart. It advances **within** a
            chunk as well as between chunks — a chunk checkpoints every
            `batch_size` rows — so it keeps moving during the minutes a commit
            chunk takes, and is the field to watch to tell a slow run from a
            stopped one.
          minimum: 0
          title: Processed Rows
          type: integer
        poll_url:
          description: Path to poll for this job's status and result.
          title: Poll Url
          type: string
        started_at:
          anyOf:
            - format: date-time
              type: string
            - type: 'null'
          description: When the run began. Null before it starts.
          title: Started At
        completed_at:
          anyOf:
            - format: date-time
              type: string
            - type: 'null'
          description: When the run reached a terminal state.
          title: Completed At
        error:
          anyOf:
            - type: string
            - type: 'null'
          description: Why the run failed. Null unless `status` is `failed`.
          title: Error
        result:
          additionalProperties: true
          description: >-
            The engine's census. Empty until the first chunk finishes, then the
            running total for the chunks merged so far, and the whole feed's
            figures once `status` is `completed` — so read `status`, not this
            field, to decide whether a run is done. Carries `rows_read`,
            `products_created` / `products_enriched` / `products_unchanged`,
            `listings_created` / `listings_updated`, `dual_listed` (products
            another retailer's feed had already created — the cross-retailer
            overlap), `duplicate_gtins_in_feed`, `quarantined` with
            `quarantine_reasons`, brand and ingredient counts, `field_decisions`
            (per field, how each value was decided), `reclaimable_fields`, and
            `quarantine` — the refused rows themselves, capped at 500 entries
            with `quarantined` remaining the true total. `errors` is capped the
            same way, with `errors_total` as its true count. `census` is the
            same figures as human-readable lines.
          title: Result
          type: object
      required:
        - job_id
        - status
        - source
        - dry_run
        - reclaim_unattributed
        - staged_rows
        - chunk_rows
        - chunks_total
        - chunks_done
        - processed_rows
        - poll_url
      title: CatalogJobOut
      type: object
    ErrorOut:
      description: |-
        RFC 9457 Problem Details response.

        All API errors are returned in this format with Content-Type:
        application/problem+json.
      examples:
        - detail: The requested resource was not found.
          error_code: not_found
          retryable: false
          status: 404
          timestamp: '2026-03-31T12:00:00+00:00'
          title: Not Found
          type: https://closient.com/docs/errors/not_found
        - detail: Validation error.
          details:
            - loc:
                - body
                - name
              msg: Field required
              type: missing
          error_code: validation_error
          retryable: false
          status: 422
          timestamp: '2026-03-31T12:00:00+00:00'
          title: Validation Error
          type: https://closient.com/docs/errors/validation_error
        - detail: Rate limit exceeded. Please try again later.
          error_code: rate_limited
          retry_after: 31
          retryable: true
          status: 429
          timestamp: '2026-03-31T12:00:00+00:00'
          title: Rate Limited
          type: https://closient.com/docs/errors/rate_limited
      properties:
        type:
          description: URI reference identifying the error type.
          title: Type
          type: string
        title:
          description: Short human-readable summary of the error.
          title: Title
          type: string
        status:
          description: HTTP status code.
          title: Status
          type: integer
        detail:
          description: Human-readable explanation of this specific occurrence.
          title: Detail
          type: string
        error_code:
          description: Machine-readable error code (e.g. not_found, unauthorized).
          title: Error Code
          type: string
        retryable:
          default: false
          description: Whether retrying the same request can succeed.
          title: Retryable
          type: boolean
        timestamp:
          description: ISO 8601 timestamp of when the error occurred.
          title: Timestamp
          type: string
        retry_after:
          anyOf:
            - type: integer
            - type: 'null'
          description: Seconds to wait before retrying (when applicable).
          title: Retry After
        owner_action_required:
          anyOf:
            - type: boolean
            - type: 'null'
          description: Whether the error requires account owner intervention.
          title: Owner Action Required
        details:
          description: Additional context (validation errors, etc.).
          title: Details
      required:
        - type
        - title
        - status
        - detail
        - error_code
        - timestamp
      title: ErrorOut
      type: object
    CatalogImportSourceEnum:
      description: |-
        Retailer catalog feeds the import engine can merge (C-5966).

        Mirrors the keys of
        ``apps.products.services.retailer_catalog_import.RETAILER_BY_SOURCE``,
        which are themselves ``ProductDataSource`` values.

        Published as an enum rather than a free string because the value is not
        a label: it selects the canonical retailer, that retailer's storefront,
        and the measured per-field confidence table that decides every contested
        field on an overlapping GTIN. An unregistered value has no retailer to
        attach to, so the engine would raise ``KeyError`` — an enum makes that a
        schema-level rejection naming the valid options instead of a 422 the
        published contract did not predict.
      enum:
        - ulta_catalog
        - sephora_catalog
      title: CatalogImportSourceEnum
      type: string
  securitySchemes:
    APIKeyHeaderAuth:
      type: apiKey
      in: header
      name: X-API-Key
    OAuthTokenAuth:
      type: http
      scheme: bearer
    CookieGatedSessionAuth:
      type: apiKey
      in: cookie
      name: sessionid

````