file Command Explained
file answers a question that seems like it should have an obvious answer but often does not: what kind of file is this, actually? Extensions lie, get stripped, or were never set correctly in the first place. file looks past the name entirely and inspects the real content.
Basic usage
file document.pdf
# document.pdf: PDF document, version 1.7
file photo.jpg
# photo.jpg: JPEG image data, JFIF standard 1.01
file script.sh
# script.sh: Bourne-Again shell script, ASCII text executable
For a well-behaved file with a correct extension, this might seem redundant. Its real value shows up the moment a file’s name and its actual content disagree.
How file actually identifies content: magic bytes
head -c 4 photo.jpg | xxd
# 0000000: ffd8 ffe0
Most file formats begin with a distinctive, fixed sequence of bytes at the very start of the file, often called magic bytes or a magic number. A JPEG always starts with FF D8 FF. A PDF always starts with the literal text %PDF. A gzip-compressed file always starts with 1F 8B. file maintains a large database of these known signatures and checks the beginning of the target file against them to determine what it actually is, entirely independent of whatever the filename claims.
# See the raw database file's rules can be inspected on many systems
file --version
locate magic.mgc 2>/dev/null
When the extension lies
mv photo.jpg photo.txt
file photo.txt
# photo.txt: JPEG image data, JFIF standard 1.01
Renaming a file does not change what is actually inside it. file catches this instantly, which is genuinely useful in a few common real-world situations: verifying a suspicious download really is what it claims to be, diagnosing a broken file transfer that silently corrupted an extension, or investigating a file that arrived with no extension at all and no other clue about its format.
file mystery_download
# mystery_download: gzip compressed data, from Unix
Text files: ASCII vs UTF-8 vs data
file readme.txt
# readme.txt: ASCII text
file readme_international.txt
# readme_international.txt: UTF-8 Unicode text
file program.bin
# program.bin: data
file distinguishes between plain ASCII text (every byte within the basic 128-character range), UTF-8 encoded text (which supports a much wider range of characters while remaining backward-compatible with ASCII), and generic binary “data,” which is what file reports when it cannot identify any recognizable structure or known format at all.
Identifying executables and scripts
file /bin/ls
# /bin/ls: ELF 64-bit LSB pie executable, x86-64
file deploy.sh
# deploy.sh: Bourne-Again shell script, ASCII text executable
file setup.py
# setup.py: Python script, ASCII text executable
For compiled binaries, file reports the executable format (ELF is the standard Linux executable format), the target architecture, and whether it is statically or dynamically linked. For scripts, it typically reads the shebang line (the #!/bin/bash or #!/usr/bin/env python3 at the very top) to identify which interpreter the script expects.
Checking archives and compressed files
file backup.tar.gz
# backup.tar.gz: gzip compressed data, was "backup.tar"
file archive.zip
# archive.zip: Zip archive data, at least v2.0 to extract
file corrupted.zip
# corrupted.zip: data
A file reported as generic “data” instead of a recognizable format, when you expected it to be a valid archive, is a strong signal that the download or transfer was incomplete or corrupted, since a valid archive of that type would normally begin with its expected magic bytes.
Using file across many files at once
file *
# script.sh: Bourne-Again shell script, ASCII text executable
# photo.jpg: JPEG image data, JFIF standard 1.01
# notes.txt: UTF-8 Unicode text
# archive.zip: Zip archive data, at least v2.0 to extract
file -i *
# same idea, but reports MIME types instead of a human description
# script.sh: text/x-shellscript; charset=us-ascii
file -i (or --mime) reports a MIME type instead of the default human-readable description, which is useful when scripting around file’s output, since MIME types are a more standardized, machine-parseable format than the free-text descriptions file normally produces.
Frequently Asked Questions
What does the file command do?
file examines the actual content of a file and reports what type of file it really is, such as a JPEG image, a PDF document, an ELF executable, a shell script, or plain text. It does this by inspecting the file’s content, not by trusting its filename or extension, which makes it reliable even when a file has been renamed or has no extension at all.
How does file determine a file’s type without relying on its extension?
Most file formats begin with a distinctive sequence of bytes called a magic number or magic bytes, a fixed pattern that identifies the format. file maintains a database of these known signatures (often at /usr/share/file/magic or similar) and compares the beginning of the target file against them. A JPEG always starts with the same few bytes, a PDF always starts with %PDF, and so on, regardless of what the file is named.
Why would a file’s extension not match what file reports?
File extensions are just naming conventions with no enforcement behind them; nothing stops a file from being renamed to end in .txt while actually containing JPEG image data, or a script being saved without any extension at all. This can happen accidentally (a bad download, a broken conversion) or deliberately, and file exposes the mismatch by reporting what the content actually is rather than trusting the name.
What does it mean when file reports “ASCII text” vs “UTF-8 Unicode text”?
Both indicate the file contains plain, human-readable text rather than binary data, but they describe the character encoding used. ASCII text means every byte in the file falls within the basic 128-character ASCII range. UTF-8 Unicode text means the file uses UTF-8 encoding, which can represent a much broader range of characters, including accented letters, symbols, and non-Latin scripts, and is backward-compatible with ASCII for the characters both encodings share.
How do I check if a file is a script and which interpreter it expects?
Run file on the script, and if the file starts with a shebang line (like #!/bin/bash), file typically reports it directly, for example as “Bourne-Again shell script, ASCII text executable.” This tells you which interpreter the script was written to run under without needing to open the file and read the first line yourself.
Can file identify compressed or archived files like .tar.gz or .zip?
Yes, file recognizes the magic bytes of common compression and archive formats and reports them accordingly, such as “gzip compressed data” or “Zip archive data,” often including additional details like the compression method or whether the archive is empty. This is useful for confirming that a downloaded archive is actually valid and complete before attempting to extract it.