Key-based deduplication

Define what duplicate means for this dataset.

Data workspace

Ready in this tab

Source

0 B Processed locally
Import settingsAuto-detect enabled
JSON text

Overrides apply on the next analysis. File decoding and every transform run only in this tab.

Paste data or choose a file to begin.

Analysis results

Analysis canvas

Ready when you are.

  1. Add a local file or paste data.
  2. Analyze format, quality, and likely sensitive fields.
  3. Review, clean, and export the in-tab working copy.
01 / Parse

Know the shape of the data.

Detect delimiters, validate structured text, and surface useful error locations.

02 / Review

Spot quality and privacy risks.

Measure completeness, duplicates, mixed values, and likely identifiers before sharing.

03 / Prepare

Shape a working copy locally.

Filter, sort, select, rename, deduplicate by keys, validate schemas, and export without changing the original file.

Focused utilities

Start with the job in front of you.

Every route opens the same private engine with guidance tailored to that format.

Practical guide

Exact rows and duplicate identities are different problems

Two records can represent the same entity even when timestamps or notes differ. Key-based deduplication lets you state which columns establish identity and which occurrence should survive.

  • 01Profile exact duplicate rows before changing the dataset.
  • 02Choose one or more columns as the identity key.
  • 03Keep the first or last matching record consistently.
  • 04Undo the operation or reset to the original analyzed rows.
Does Parse The Data upload my files?

No. Parsing, profiling, cleanup, and export run inside your browser tab. The application does not send dataset contents to our server or any third party.

What file types are supported?

CSV, TSV, JSON, and JSON Lines (JSONL or NDJSON) are supported as pasted text or local files up to 10 MB. CSV files can also be decoded as UTF-8, UTF-16, or Windows-1252.

Can it detect every privacy risk?

No. The privacy scan uses deterministic patterns and column-name hints. It can find common emails, phone numbers, IP addresses, SSNs, and payment-card patterns, but it is not a substitute for a formal privacy review.

Designed for restraint

The server delivers the tools. Your dataset stays in the browser.

The app files arrive first. Parsing, profiling, and cleanup then happen in a browser worker on your device. The dataset has no upload path.

  • No accounts
  • No cloud storage
  • No third-party scripts

Use the PII scan as an early warning—not as legal advice or a compliance certification.