Extract emails from a PDF, Word or CSV file
By the end you will be able to turn a PDF, Word file or spreadsheet into a clean, de-duplicated list of addresses filtered by domain, and know what to check before you use it.
To pull every email address out of a document, open the email extractor, drop in your PDF, Word, Excel or CSV file (or paste the text) and the list appears instantly, with repeats removed and ready to copy or export. The file is read inside your browser and never sent to a server, so you can run it on your real file right now, even if it contains other people's addresses.
How do you extract emails from a PDF, Word or Excel file?
The steps are the same for every format and take a few seconds. There is nothing to install and no account to create.
- Load your material. Drop one or more files into the box, or paste text, HTML or code into the text area. You can mix both: the result combines everything.
- Set the options. Remove duplicates and Ignore case are on by default. This is also where you choose which domains to keep or exclude.
- Check the list. You will see the number of unique addresses and the number of duplicates dropped. If you loaded several files, each one shows how many addresses it contributed.
- Copy or export. Copy to the clipboard, or download as TXT (one address per line), CSV (a single column headed
email) or JSON (an array of strings).
If you want the result organized, switch on Group by domain: addresses are gathered under each domain, most frequent first. It is a quick way to see which companies appear most in a list.
Which file types it reads, and which it doesn't
It reads plain text and the everyday formats: .txt, .csv, .tsv, .json, .html, .xml, .vcf (contact cards), .eml (saved emails) and .md, plus .pdf, .docx and .xlsx. The limit is 50 MB per file.
Three limits are worth knowing. First, a scanned PDF is only an image with no text in it, and this tool does not do OCR. In that case run the file through an OCR for PDFs first, or follow the guide to extract text from a scanned PDF. Second, a password-protected PDF will not open: the tool tells you so, and you need to remove the PDF password first (you must know it; see the guide to removing a PDF password). Third, the old .doc and .xls formats are not supported, so save them as .docx or .xlsx.
Two practical details. From a Word file it also reads headers, footers, notes and comments, and it picks up addresses hidden behind a mailto: hyperlink even when the visible text says something else, such as "contact us". In Excel it scans every sheet, not just the first. If your list is a CSV with garbled accents, the tool tries UTF-8 and falls back to Windows-1252, which is what Excel often writes. To turn that CSV into a spreadsheet, use CSV to Excel.
Duplicates, case and domain filters
Real lists are rarely clean: the same address shows up in every signature, in the header and in the body, sometimes with different capitalization. With Remove duplicates and Ignore case on, all those variants collapse into one. A technical aside: the part before the @ could in theory be case-sensitive, but in practice mail providers treat it as case-insensitive, which is why normalizing is safe. If you need the original spelling, switch the option off.
Domain filters solve the other usual problem, noise. Exclude your own domain (your company signature is everywhere), keep only company.com to see which people from one client joined a thread, or separate personal addresses (gmail.com, outlook.com) from corporate ones. A domain filter also covers its subdomains.
The tool also drops things that look like emails but are not. Pasting HTML or CSS turns up asset names such as logo@2x.png, which match the shape of an address; they are discarded when they end in common image, font or script extensions.
Obfuscated addresses: when to turn on de-obfuscation
Many sites and documents write addresses so bots cannot read them, like name [at] domain [dot] com or name(at)domain(dot)com. The option is off by default. When you turn it on, the tool understands at and dot (and the Spanish arroba and punto) inside brackets, parentheses, braces or angle brackets, plus the HTML encodings of the at sign (@, @) and of the dot.
It does not replace a bare " at ", because that would turn ordinary sentences like "meet at noon" into fake addresses. It also cannot recover addresses built by JavaScript or stored inside an image. And because addresses with non-ASCII characters (a domain with an accented letter, for example) are not recognized, always review the list before using it: no automatic extraction is perfect.
Is it legal to use the addresses you extract?
Extracting an address and having the right to write to it are different things, and the tool does not check the second. What follows is general orientation, not legal advice.
- GDPR (European Union): an email address linked to a person is personal data, work addresses included (
first.last@company.com). Processing it usually needs a lawful basis, such as consent or a well-justified legitimate interest. - LFPDPPP (Mexico): governs how private individuals and companies handle personal data, and revolves around the privacy notice and the data subject's consent.
- CAN-SPAM (United States): regulates commercial email. It requires you to identify the sender, avoid misleading subject lines, include a postal address and offer a working way to opt out.
Lower-risk uses are the familiar ones: migrating your own subscriber list, collecting the contacts from an event you organized, or auditing a document before you publish it to see whether an address slipped in that should not be there. Higher-risk ones are bulk sends to addresses nobody gave you. If you are unsure about a commercial send, ask a professional.
Why not upload a contact list to an unknown website?
Because a list of emails is other people's personal data, and the moment you upload it to someone else's server you lose control of it: you do not know whether it is stored, for how long or who can see it. If the list belongs to your customers, you may also be breaking what you promised them.
Here the file is processed in your browser and there is no upload. You do not have to take that on trust: open your browser's developer tools (F12), go to the Network tab, drop in a file and confirm that no request carries its contents. Run the same test on any other site before you hand it a sensitive file.
Frequently asked questions
How do I extract email addresses from a PDF?
Open the extractor, drop the PDF into it and the list appears on its own. It works on PDFs with selectable text. A scanned PDF is just an image with no text to read, so it needs OCR first. If the PDF is password-protected, remove the password before opening it.
Can I extract emails from a Word document or an Excel spreadsheet?
Yes, as long as they are .docx or .xlsx files. From Word it reads the body, headers, footers, notes and comments, plus mailto links. From Excel it reads every sheet and email hyperlinks. The older .doc and .xls formats are not supported, so save them as .docx or .xlsx first.
How do I remove duplicate email addresses?
Leave Remove duplicates switched on, which is the default. Each address appears once, and the counter shows how many repeats were dropped. With Ignore case on, Ana@Company.com and ana@company.com count as the same address.
Can I keep only one domain, or leave one out?
Yes. There are two filter fields: one to keep only certain domains and one to exclude domains. Filters include subdomains, so excluding company.com also drops name@mail.company.com. You can list several domains separated by commas.
Is it legal to email the addresses I extract?
It depends on whose addresses they are, how you got them and where you are. Rules such as the GDPR in the EU, Mexico's LFPDPPP or the US CAN-SPAM Act govern personal data and commercial email, and being able to extract an address does not give you permission to write to it. This is general information, not legal advice: for any commercial use, talk to a professional.
Is it safe to use an online extractor on a contact list?
Only if the file never leaves your device. This tool reads the file in your browser and does not upload it to any server, and you can check that in your browser's Network tab. If you cannot verify a site works that way, do not upload other people's contact lists to it.
Pull the emails out of your document now
PDF, Word, Excel or text, 100% in your browser. Free, no signup.
Open the email extractor →Related tools
- Extract text from a PDF: pull the full text out to review before searching for addresses.
- OCR PDF: give scanned PDFs a text layer so their emails can be extracted.
- Remove PDF password: open a protected PDF (you need the password) so it can be read.
- CSV to Excel: move your exported list into a spreadsheet.
You might also like: convert CSV to Excel without breaking your data.