Public API

The exported surface of the package. Everything here is covered by semantic versioning; see Internals for the non-exported helpers.

MissingPatterns.MissingPatterns — Module
MissingPatterns

Terminal-based text visualizations for missing data patterns in any Tables.jl-compatible source (DataFrames, CSV.File, NamedTuples of vectors, XLSX tables, ...), with zero plotting-library dependencies.

Exported API:

  • plotmissing — where/how much is missing (heatmap; optional temporal grouping via by/period).
  • missingpatterns — which columns are missing together (unique row patterns, à la mice::md.pattern()).
  • missingcooccurrence — pairwise ϕ/Jaccard association of missingness masks.
  • missingsummary — per-column table with counts, % and a sparkline of where along the rows the missing values concentrate.
  • plotmissingdiff — before/after comparison (e.g. auditing an imputation step).
  • missinghtml — the heatmap as a standalone HTML fragment.
source
MissingPatterns.missingcooccurrence — Method
missingcooccurrence([io::IO=stdout], tbl; method=:phi, cell_chars=5,
                     name_width=4, color=:auto, missing_color="#f3a9a9",
                     max_cols=20, isna=ismissing)

Display the pairwise co-occurrence matrix of missingness between columns — ϕ coefficient (default) or Jaccard index of the missing masks. High positive values mean "these columns go missing together", which is strong evidence against MCAR and directly informs imputation strategy.

If the table has more than max_cols columns, the max_cols columns with the most missing values are shown (they carry the information) and a note reports how many were omitted. Diagonal cells are —; degenerate pairs (columns with no missing values) are ·.

Cell text is the coefficient (-1.00…1.00); with color enabled, cell intensity scales with |value| using missing_color.

  • isna: predicate deciding what counts as an absent value (default: ismissing), as in plotmissing.
source
MissingPatterns.missingdrop — Method
missingdrop([io::IO=stdout], tbl; bar_width=30, color=:auto,
             missing_color="#f3a9a9", isna=ismissing)

Show what dropping columns buys complete-case analysis: at each step the column whose removal completes the most rows, and how many rows survive dropmissing once it is gone.

Where missingrows prices listwise deletion for the table as it stands, this prices the alternative — trading a variable for rows — and names the variable worth trading. The cells column (complete × columns left) sizes the surviving complete-case block; the step that maximizes it is flagged, since past that point each drop costs more in columns than it returns in rows.

Arguments

  • bar_width::Int: width in characters of the longest bar (default: 30).
  • color::Symbol / missing_color::String: as in plotmissing.
  • isna: predicate deciding what counts as absent (see plotmissing).

Returns nothing; use missingdropstats for the same numbers as data.

source
MissingPatterns.missingdropstats — Method
missingdropstats(tbl; isna=ismissing) -> Vector{<:NamedTuple}

The listwise-deletion trade-off as a Tables.jl-compatible row table — the data behind missingdrop.

Complete-case analysis discards every row with a missing value, and a single sparse column can be responsible for most of that loss. This walks the greedy path of column drops — at each step removing the column that turns the most rows complete — and reports what complete-case analysis looks like after each one, so the cost of keeping a column is expressed in the rows it costs.

One row per step, starting from the untouched table, with fields:

  • ndropped::Int — columns dropped so far (0 on the first row).
  • dropped::Union{Nothing,Symbol} — the column dropped at this step; nothing on the first row, which is the table as given.
  • ncols::Int — columns left.
  • complete::Int — rows with no missing value among the columns left, i.e. what dropmissing would keep.
  • pct::Float64 — 100 * complete / nrows.
  • cells::Int — complete * ncols, the size of the complete-case block that survives. It rises while a drop buys more rows than it costs columns and falls afterwards, so its maximum is the natural stopping point.

The walk ends once every row is complete or one column is left. isna is the "counts as absent" predicate (see plotmissing).

Examples

using DataFrames
df = DataFrame(a=[1, 2, missing, 4], b=[1, missing, missing, 4], c=[1, 2, 3, 4])

DataFrame(missingdropstats(df))

# the drop that leaves the largest complete-case block
steps = missingdropstats(df)
best = steps[argmax([r.cells for r in steps])]

See also missingrowstats for the distribution this optimizes against.

source
MissingPatterns.missinghtml — Method
missinghtml(tbl; max_rows=200, max_cols=60, missing_color="#f3a9a9",
            emphasis=:present, title="Missing data",
            by=nothing, period=nothing, isna=ismissing,
            order=:table) -> String
missinghtml(path::AbstractString, tbl; kwargs...) -> path

Render the missing-data heatmap as a standalone HTML fragment (dark-themed <div>, no external CSS/JS) — suitable for pasting into a blog post, notebook export, or report. Column headers are rotated for readability; every cell carries a tooltip with its row range and exact missing percentage. The same compression engine and color ramp as plotmissing are used, so the two outputs always agree — HTML just affords a much larger grid (defaults: 200×60 blocks).

by/period group rows by a column's values instead of by position, exactly as in plotmissing, so a grouped report reads the same in both media. isna and order likewise behave exactly as they do there.

The one-argument form returns the HTML String; the two-argument form writes it to path and returns the path.

source
MissingPatterns.missingpairstats — Method
missingpairstats(tbl; isna=ismissing) -> Vector{<:NamedTuple}

Pairwise co-occurrence of missingness as a Tables.jl-compatible row table — the data behind missingcooccurrence, without the rendering and without the max_cols display cap.

One row per unordered pair of distinct columns (ncols*(ncols-1)/2 rows, in upper-triangle order), with fields:

  • a::Symbol, b::Symbol — the two column names.
  • phi::Float64 — ϕ coefficient of the two missingness masks.
  • jaccard::Float64 — Jaccard index of the same masks.
  • n11::Int — rows where both a and b are missing.
  • n1::Int, n2::Int — rows where a (resp. b) is missing.
  • nrows::Int — row count of the table.

Both coefficients fall out of the same n11/n1/n2 counts, so both are returned rather than selected by a method keyword: the schema stays fixed regardless of which one you look at.

Undefined coefficients come back as NaN, and the two differ on when that happens: phi is NaN whenever either column is entirely missing or entirely present (the 2x2 table is degenerate), while jaccard is NaN only when neither column has a single missing value — against a column with no missing values the union is still non-empty, so Jaccard is a well-defined 0.0.

Examples

pairs = missingpairstats(tbl)

# most co-missing pairs; `first` rather than `[1:5]`, which throws on a
# table with fewer than five pairs
first(sort(pairs; by = r -> -r.phi), 5)

filter(r -> r.n11 > 0 && r.jaccard > 0.5, pairs)

See also missingstats, missingpatternstats.

source
MissingPatterns.missingpatterns — Method
missingpatterns([io::IO=stdout], tbl; max_patterns=20, cell_chars=5,
                 char_missing='█', char_present='░', name_width=4,
                 color_cells=false, missing_color="#f3a9a9",
                 emphasis=:present, show_bar=true, min_pct=0.0,
                 isna=ismissing)

Display the unique row-wise missingness patterns found in a Tables.jl-compatible tbl, sorted by descending frequency — i.e. which columns tend to be missing together.

This complements plotmissing, which shows where/how much is missing: missingpatterns shows which combinations of missing columns actually occur, the same diagnostic produced by R's mice::md.pattern(). Useful for reasoning about the missingness mechanism (e.g. if two columns are always missing together, that's rarely MCAR) and for choosing an imputation strategy. For a correlation-style view of the same question, see missingcooccurrence.

Arguments

  • io::IO: output stream (default: stdout).
  • tbl: any Tables.jl-compatible table.
  • max_patterns::Int: maximum number of patterns to display, most-frequent first (default: 20). Patterns beyond this are summarized in a trailing count.
  • cell_chars::Int: number of repeated characters per cell (default: 5, max: 80).
  • char_missing::Char: character for a missing column in a pattern (default: '█').
  • char_present::Char: character for a present column in a pattern (default: '░').
  • name_width::Int: max characters shown for column names before truncating with … (default: 4; set to 0 to show full names, bounded by cell width).
  • color_cells::Bool: apply the same color ramp as plotmissing (default: false).
  • missing_color::String: hex color of the ramp, as in plotmissing (default: "#f3a9a9").
  • emphasis::Symbol: :present (present columns colored, missing dark) or :missing (inverted), as in plotmissing (default: :present).
  • show_bar::Bool: append an UpSet-style horizontal frequency bar per pattern, scaled to the most common displayed pattern (default: true).
  • min_pct::Float64: hide patterns matching fewer than this percentage of rows (default: 0.0 — show all). Hidden patterns are reported in the trailing summary line.
  • isna: predicate deciding what counts as an absent value (default: ismissing), as in plotmissing.

Returns

  • nothing. The table is written to io.
source
MissingPatterns.missingpatternstats — Method
missingpatternstats(tbl; isna=ismissing) -> Vector{<:NamedTuple}

Unique row-wise missingness patterns as a Tables.jl-compatible row table — the data behind missingpatterns, without the rendering and without the max_patterns/min_pct display caps: every pattern is returned.

One row per distinct pattern, most frequent first (ties broken by first appearance, matching the displayed order), with fields:

  • pattern::NamedTuple — one Bool per column, keyed by column name; true means missing in this pattern.
  • nmissing::Int — how many columns are missing in the pattern.
  • n::Int — rows matching it.
  • pct::Float64 — 100 * n / nrows.

Examples

tbl = (age = [34, missing, 51], income = [missing, 4200, 5100])
ps = missingpatternstats(tbl)

ps[1].pattern.age                 # was `age` missing in the most frequent pattern?
filter(r -> r.nmissing == 0, ps)  # the fully-complete pattern, if any
Note

pattern is a NamedTuple with one field per column, so its type depends on the table's width. On tables with many hundreds of columns this costs noticeable compile time on first call; the counts themselves stay cheap, since they come from the deduplicated pattern table.

See also missingstats, missingpairstats.

source
MissingPatterns.missingreport — Method
missingreport(tbl; kwargs...) -> MissingReport

Wrap tbl in an object that renders itself as the missing-value heatmap in whichever medium displays it:

  • MIME"text/plain" (REPL, logs, files) → the terminal heatmap, identical to plotmissing.
  • MIME"text/html" (Jupyter, Pluto, Documenter, any HTML-aware display) → the HTML heatmap, identical to missinghtml.

So in a notebook missingreport(df) shows the colored HTML grid with per-cell tooltips, and the very same expression in a terminal shows the Unicode grid — with no if on the caller's side.

Keyword arguments are those of plotmissing and missinghtml; each is forwarded only to the renderer that accepts it, so per-medium defaults (e.g. a 200×60 HTML grid vs a 50×20 terminal grid) are preserved unless you override them explicitly. isna and order reach both renderers, so a report reads the same in either medium. An unknown keyword is an error at construction time, not at display time.

Examples

missingreport(df)
missingreport(df; emphasis=:missing, missing_color="#ff6600")
missingreport(df; by=:region)                 # grouped in both media
missingreport(df; layout=:compact, title="Cohort A")

# force one medium explicitly
show(stdout, MIME"text/html"(), missingreport(df))

plotmissing and missinghtml are unchanged and remain the direct, single-medium entry points.

source
MissingPatterns.missingrows — Method
missingrows([io::IO=stdout], tbl; sortby=:nmissing, bar_width=30,
             color=:auto, missing_color="#f3a9a9", isna=ismissing)

Distribution of how many values are missing per row: how many rows are complete, how many are missing exactly one value, two, and so on.

This is the view missingsummary (per column) and missingpatterns (per combination) leave out, and it is the one that answers "what would listwise deletion cost me?" — the 0 line is the complete-case count, everything below it is what dropmissing would discard.

Arguments

  • sortby::Symbol: :nmissing (ascending missing-count, default) or :rows (descending row count — most common shape first).
  • bar_width::Int: width in characters of the longest bar (default: 30).
  • color::Symbol / missing_color::String: as in plotmissing; bars are tinted by severity (nmissing / ncols), so complete rows carry the "present" color and fully-missing rows the full missing color.
  • isna: predicate deciding what counts as an absent value (default: ismissing), as in plotmissing.

Returns nothing; use missingrowstats for the same numbers as data.

source
MissingPatterns.missingrowstats — Method
missingrowstats(tbl; isna=ismissing) -> Vector{<:NamedTuple}

Row-completeness distribution as a Tables.jl-compatible row table — the data behind missingrows.

One row per observed missing-count (counts that occur zero times are omitted), ascending, with fields:

  • nmissing::Int — number of missing values in such a row (0 = complete).
  • nrows::Int — how many rows of tbl have exactly that many.
  • pct::Float64 — 100 * nrows / total rows.

Examples

rs = missingrowstats(df)
only(r.nrows for r in rs if r.nmissing == 0)   # complete-case count
sum(r.nrows for r in rs if r.nmissing > 0)     # rows lost to listwise deletion

See also missingstats for the transposed (per-column) view.

source
MissingPatterns.missingstats — Method
missingstats(tbl; isna=ismissing) -> Vector{<:NamedTuple}

Per-column missing-data statistics as a Tables.jl-compatible row table — the data behind missingsummary, without the rendering.

One row per column of tbl, in table order, with fields:

  • column::Symbol — the column's name.
  • eltype::Type — the column's element type, Missing included (missingsummary displays Base.nonmissingtype of this).
  • nmissing::Int, npresent::Int — cell counts.
  • nrows::Int — row count of the table (same on every row; carried so a single row is self-contained after filtering).
  • pct::Float64 — 100 * nmissing / nrows, or 0.0 for an empty table.

Examples

using DataFrames
df = DataFrame(a=[1, missing, 3], b=["x", "y", "z"])

DataFrame(missingstats(df))          # straight into a DataFrame
filter(r -> r.pct > 20, missingstats(df))

See also missingpatternstats, missingpairstats, missingrowstats.

source
MissingPatterns.missingsummary — Method
missingsummary([io::IO=stdout], tbl; bins=20, sortby=:missing,
                color=:auto, missing_color="#f3a9a9", isna=ismissing)

Per-column missing-data overview: name, element type, missing count, %, and a sparkline showing where along the rows the missing values concentrate — the row axis is split into bins equal blocks and each block maps to a bar height proportional to its missing fraction. A block with even one missing value renders at least the smallest bar (same visibility guarantee as plotmissing); a block with none renders blank.

Arguments

  • bins::Int: sparkline resolution (default: 20).
  • sortby::Symbol: :missing (descending missing count, default), :name, or :none (table order).
  • color::Symbol / missing_color::String: as in plotmissing; the sparkline colors bars by their missing fraction.
  • isna: predicate deciding what counts as an absent value (default: ismissing), as in plotmissing.
source
MissingPatterns.plotmissing — Method
plotmissing([io::IO=stdout], tbl; cell_chars=5, char_missing='█', char_present='░',
            name_width=4, color_cells=false, show_row_range=false,
            max_rows=50, max_cols=20,
            layout=:auto, target_lines=28, color=:auto,
            missing_color="#f3a9a9", emphasis=:present,
            by=nothing, period=nothing, isna=ismissing, order=:table)

Display a text-based heatmap of missing value patterns in any Tables.jl-compatible source (DataFrame, CSV.File, NamedTuple of vectors, ...). When the data exceeds the display limits, multiple rows/columns are grouped into a single cell using a Unicode block-character gradient (classic layout) or an ANSI-colored half-block encoding (compact layout).

Layouts

  • :classic — the original layout: one grid row per line, 3-line header, 6-line summary. Best in a full terminal with room to scroll.
  • :compact — fits the entire plot (grid + header + summary) in at most target_lines lines, so IDE/Jupyter output cells never truncate it. With color available, each output line encodes two grid rows via '▀' (foreground = top row, background = bottom row), doubling vertical resolution; without color it falls back to the glyph gradient at one row per line. Any block containing even a single missing value is rendered in a shade distinct from "fully present", so fine holes survive compression.
  • :auto (default) — uses :classic when it fits within target_lines, :compact otherwise.

Grouping

  • by::Union{Nothing,Symbol,String}: name of a column. When set, rows are grouped by the values of that column (not by position), so the vertical axis becomes honest categories/calendar time instead of arbitrary row ranges. Rows whose by value is missing form a trailing ∅ group. Row labels are always shown in this mode. If there are more groups than fit the budget, consecutive groups (in sorted order) are merged and labeled as ranges (e.g. 2004-2005, A-C).
  • period::Union{Symbol,Nothing}: nothing (default) groups by the by column's exact value — categorical grouping, works for any sortable column (String, Symbol, Int, ...). Set to :year, :quarter, :month, :week or :day to instead group by that calendar period of a Date/DateTime by column. :week follows ISO-8601, so its label carries the ISO week-year, which at a year boundary can differ from the calendar year (2024-12-30 is 2025-W01).

Arguments

  • io::IO: output stream (default: stdout).

  • tbl: any Tables.jl-compatible table.

  • cell_chars::Int: number of repeated characters per heatmap cell (default: 5, max: 80).

  • char_missing::Char: character for fully-missing cells (default: '█').

  • char_present::Char: character for fully-present cells (default: '░').

  • name_width::Int: max characters shown for column names before truncating with … (default: 4; set to 0 to show full names, bounded by cell width).

  • color_cells::Bool: apply the color ramp to classic-layout glyphs. The compact half-block layout always colors its cells (default: false).

  • show_row_range::Bool: display row-range (or period) labels in a left-hand column (default: false; forced true when by is set).

  • max_rows::Int: maximum display rows before compression in the classic layout (default: 50). Ignored by :compact, which derives its own limit from target_lines.

  • max_cols::Int: maximum display columns before compression (default: 20).

  • layout::Symbol: :auto, :classic, or :compact (default: :auto).

  • target_lines::Int: total line budget for the compact layout, including borders, header and summary (default: 28 — safely under typical IDE output-cell limits of ~30 lines).

  • color::Symbol: :auto (respect io's :color property / TTY detection), :always (force ANSI codes — use this in VS Code/Jupyter notebooks, whose output cells render ANSI but whose stdout is not a TTY), or :never (plain text — use when redirecting to a file).

  • missing_color::String: hex color ("#rrggbb") of the ramp (default: "#f3a9a9").

  • emphasis::Symbol: which side of the data carries the ink (default: :present). With :present, present data is painted in missing_color and missing data fades to dark gray — holes read as dark gaps in a colored field. With :missing, the ramp is inverted. In both modes, any block containing even one missing value renders in a shade visibly different from a fully-present block.

  • isna: what counts as an absent value (default: ismissing). Microdata often codes absence as a sentinel rather than as missing — 9/99 for "ignored", "" for a blank field — and this lets those count as holes without rewriting the table. Either a predicate applied to every column:

    plotmissing(df; isna = x -> ismissing(x) || x == 9 || x == "")

    or, better, a NamedTuple/AbstractDict of per-column predicates, with ismissing assumed for any column left out:

    plotmissing(df; isna = (criterio = x -> ismissing(x) || x == 9,
                            cid      = x -> ismissing(x) || x == ""))

    Prefer the per-column form. A sentinel belongs to a variable, not to a table: 9 means "ignored" in a coded field but is a perfectly good age, and a blanket predicate would punch holes in every column that happens to hold the value. Naming a column the table does not have is an error, not a silently ignored entry.

    In either form, test ismissing first and let || short-circuit: missing == 9 is missing, not false, and would be an error in a boolean context. The predicate applies to every count the package makes, the by column included, so a sentinel there forms the ∅ group.

  • order::Symbol: column order (default: :table). :table keeps the table's own order; :missing puts the emptiest columns first; :name sorts alphabetically; :cluster places columns that go missing together side by side, which is usually what makes a block structure visible at all — table order scatters them. :cluster seriates the ϕ matrix and so costs one extra pass over the data; columns with no missing values carry no pattern and are appended at the end. Reordering is purely a display concern and never changes a number. When columns are compressed, a reordered group is labeled by its endpoint names rather than by positional indices, which would otherwise refer to display slots.

Returns

  • nothing. The plot is written to io.
source
MissingPatterns.plotmissingdiff — Method
plotmissingdiff([io::IO=stdout], before, after; cell_chars=5, name_width=4,
                 max_cols=20, target_lines=28, color=:auto,
                 missing_color="#f3a9a9", filled_color="#a9f3c1",
                 isna=ismissing)

Compare the missingness of two same-shaped tables — typically the same dataset before and after an imputation step, or two releases of a periodic microdata file. Rendered in the compact layout (fits target_lines):

  • neutral dark gray — block unchanged;
  • tint of missing_color — block got more missing (introduced holes);
  • tint of filled_color — block got less missing (holes resolved).

With color, half-blocks encode two row blocks per line (as in plotmissing); without color, cells fall back to + (more missing), - (fewer) and · (unchanged) glyphs. The summary line reports exact cell-level counts of resolved and introduced missing values, computed by a row-aligned pass (not from block averages).

Both tables must have identical dimensions and column names, in order.

  • isna: predicate deciding what counts as an absent value (default: ismissing), as in plotmissing.
source