Locale and Internationalization on Linux

Locale and Internationalization on Linux

A locale defines how a system formats and compares text: language, character encoding, date format, decimal separator, currency, and sort order.

Most of the time it works invisibly. When it does not, the symptom is usually a script producing different output on a different machine.

What is set

locale
LANG=en_GB.UTF-8
LC_CTYPE="en_GB.UTF-8"
LC_NUMERIC="en_GB.UTF-8"
LC_TIME="en_GB.UTF-8"
LC_COLLATE="en_GB.UTF-8"
LC_MONETARY="en_GB.UTF-8"
LC_MESSAGES="en_GB.UTF-8"
LC_ALL=

The categories:

VariableControls
LC_CTYPECharacter classification, upper and lower case
LC_COLLATESort order
LC_TIMEDate and time format
LC_NUMERICDecimal separator, thousands grouping
LC_MONETARYCurrency formatting
LC_MESSAGESLanguage of program messages

Precedence, lowest to highest:

  1. LANG, the default for everything
  2. Individual LC_* variables, overriding LANG for one category
  3. LC_ALL, overriding all of them

LC_ALL is deliberately blunt. Setting it in a shell profile is a mistake, because it removes the ability to override any single category. Setting it in a script is exactly right, for the opposite reason.

A common desktop arrangement:

export LANG=en_GB.UTF-8
export LC_TIME=en_GB.UTF-8      # day/month/year
export LC_NUMERIC=en_US.UTF-8   # decimal point, not comma

The one that breaks scripts

printf '%s\n' apple Banana cherry Date | sort

Under C:

Banana
Date
apple
cherry

Under en_US.UTF-8:

apple
Banana
cherry
Date

C sorts by byte value, so all uppercase precedes all lowercase. en_US.UTF-8 uses dictionary collation, case-insensitive and ignoring punctuation.

Both are correct for their purpose. The second is what a human wants in a list. The first is what a program needs when it will parse the output, and it is the only one that is the same everywhere.

This affects more than sort:

# ranges in grep and bracket expressions
echo "Hello" | grep '[a-z]'        # matches under en_US, not under C
ls [A-Z]*                          # different files matched

# comm and join require consistently sorted input
sort file1 > a; sort file2 > b; comm a b    # broken if locales differ

comm and join silently produce wrong output when their inputs were sorted under a different collation. The failure is quiet, which makes it worse.

Set LC_ALL=C in scripts:

#!/usr/bin/env bash
export LC_ALL=C
set -euo pipefail

Deterministic behaviour regardless of the machine. Our bash error handling guide covers the rest of that preamble.

Where you need locale-aware sorting for a human, ask for it explicitly:

LC_COLLATE=en_US.UTF-8 sort names.txt

Generating locales

locale -a           # what exists
locale -a | grep en # filter

Requesting a locale that is not generated gives you a warning and a silent fall back to C:

bash: warning: setlocale: LC_ALL: cannot change locale (de_DE.UTF-8)

Debian and Ubuntu:

sudo nano /etc/locale.gen       # uncomment the lines you want
sudo locale-gen
sudo update-locale LANG=en_GB.UTF-8

Or non-interactively:

sudo sed -i 's/^# *\(en_GB.UTF-8\)/\1/' /etc/locale.gen
sudo locale-gen

Fedora and RHEL use langpacks:

sudo dnf install glibc-langpack-en glibc-langpack-de

Minimal container images frequently ship with no locales at all, which is why LC_ALL=C.UTF-8 is the sensible default in a Dockerfile:

ENV LANG=C.UTF-8 LC_ALL=C.UTF-8

C.UTF-8

Worth knowing about. It handles UTF-8 correctly, so multibyte characters work, while keeping byte-order collation and untranslated messages.

That combination is usually what you want on a server: predictable sorting, English error messages that match what is in the documentation, and no mangling of non-ASCII filenames.

export LANG=C.UTF-8

Plain C without the encoding treats input as single-byte, which corrupts UTF-8 filenames and output. Prefer C.UTF-8 over bare C for anything that touches real text.

The SSH warning

-bash: warning: setlocale: LC_ALL: cannot change locale (en_GB.UTF-8)

Your client is forwarding its locale to a server that does not have it generated. Two fixes.

Generate it on the server, or stop the client sending it:

# ~/.ssh/config
Host *
    SendEnv -LANG -LC_*

Server side, the corresponding setting is AcceptEnv in sshd_config. Our SSH config builder covers the client file.

Time and numbers

LC_TIME=en_US.UTF-8 date     # Sun Sep 14 10:30:00 UTC 2026
LC_TIME=de_DE.UTF-8 date     # So 14 Sep 2026 10:30:00 UTC
LC_TIME=C date               # Sun Sep 14 10:30:00 UTC 2026

For anything machine-readable, bypass locale entirely:

date -u +%Y-%m-%dT%H:%M:%SZ
date +%s

ISO 8601 and Unix timestamps are unambiguous and sort correctly as strings. Never parse localised date output.

Numeric formatting catches people in a different way:

LC_NUMERIC=de_DE.UTF-8 printf '%.2f\n' 3.14
# 3,14

A comma decimal separator breaks anything consuming the output as a number. This is a real source of bugs in scripts that do arithmetic and then feed the result to another program.

LC_ALL=C awk '{sum += $1} END {print sum}' numbers.txt

Filenames

Filenames are bytes, not text. The kernel does not know or care about encoding.

Practical consequences: a filename created under one encoding displays as nonsense under another, and ls may show question marks for bytes it cannot interpret.

convmv -f iso-8859-1 -t utf-8 --notest -r /path/to/files
ls | cat -v      # show non-printing bytes literally

UTF-8 everywhere avoids this entirely, which is the state of most systems now.

Frequently Asked Questions

What is the difference between LANG, LC_ALL, and the individual LC_ variables?

LANG is the default for every category. Individual variables such as LC_TIME and LC_NUMERIC override it for that category only. LC_ALL overrides everything including the individual variables, which is why it is the right tool for scripts and the wrong one for a user profile.

Why does my script behave differently on another machine?

Almost always locale. Sort order, decimal separators, and date formats all change with the locale, so a script that works in one and produces wrong output in another is extremely common. Setting LC_ALL=C at the top of the script makes behaviour deterministic everywhere.

How do I add a locale that is not available?

On Debian and Ubuntu, uncomment the line in /etc/locale.gen and run locale-gen. On Fedora and RHEL, install the relevant langpacks package. Running locale -a shows what is currently generated, and requesting one that is not generated produces a warning and a silent fallback to C.

Why do I see errors about setting locale when I log in over SSH?

Your SSH client is forwarding locale variables the server does not have generated. Either generate the locale on the server, or stop the client sending them by removing SendEnv LANG LC_ from your SSH configuration.

Should I use en_US.UTF-8 or C.UTF-8?

C.UTF-8 is a good default for servers and containers because it handles UTF-8 correctly while keeping byte sort order and unlocalised messages, which makes behaviour predictable. Use a full locale on a desktop where you want correctly formatted dates, numbers, and collation in your own language.

Why does sort produce different results with different locales?

Collation rules are locale-specific. Under C, sorting is by byte value so all uppercase letters precede all lowercase ones. Under en_US.UTF-8, sorting is case-insensitive and ignores punctuation, which is correct for human-readable lists and wrong for anything a program will parse.