Public API
The exported surface of the package. Everything here is covered by semantic versioning; see Internals for the non-exported helpers.
MissingPatterns.MissingPatternsMissingPatterns.missingcooccurrenceMissingPatterns.missingdropMissingPatterns.missingdropstatsMissingPatterns.missinghtmlMissingPatterns.missingpairstatsMissingPatterns.missingpatternsMissingPatterns.missingpatternstatsMissingPatterns.missingreportMissingPatterns.missingrowsMissingPatterns.missingrowstatsMissingPatterns.missingstatsMissingPatterns.missingsummaryMissingPatterns.plotmissingMissingPatterns.plotmissingdiff
MissingPatterns.MissingPatterns — ModuleMissingPatternsTerminal-based text visualizations for missing data patterns in any Tables.jl-compatible source (DataFrames, CSV.File, NamedTuples of vectors, XLSX tables, ...), with zero plotting-library dependencies.
Exported API:
plotmissing— where/how much is missing (heatmap; optional temporal grouping viaby/period).missingpatterns— which columns are missing together (unique row patterns, à lamice::md.pattern()).missingcooccurrence— pairwise ϕ/Jaccard association of missingness masks.missingsummary— per-column table with counts, % and a sparkline of where along the rows the missing values concentrate.plotmissingdiff— before/after comparison (e.g. auditing an imputation step).missinghtml— the heatmap as a standalone HTML fragment.
MissingPatterns.missingcooccurrence — Methodmissingcooccurrence([io::IO=stdout], tbl; method=:phi, cell_chars=5,
name_width=4, color=:auto, missing_color="#f3a9a9",
max_cols=20, isna=ismissing)Display the pairwise co-occurrence matrix of missingness between columns — ϕ coefficient (default) or Jaccard index of the missing masks. High positive values mean "these columns go missing together", which is strong evidence against MCAR and directly informs imputation strategy.
If the table has more than max_cols columns, the max_cols columns with the most missing values are shown (they carry the information) and a note reports how many were omitted. Diagonal cells are —; degenerate pairs (columns with no missing values) are ·.
Cell text is the coefficient (-1.00…1.00); with color enabled, cell intensity scales with |value| using missing_color.
isna: predicate deciding what counts as an absent value (default:ismissing), as inplotmissing.
MissingPatterns.missingdrop — Methodmissingdrop([io::IO=stdout], tbl; bar_width=30, color=:auto,
missing_color="#f3a9a9", isna=ismissing)Show what dropping columns buys complete-case analysis: at each step the column whose removal completes the most rows, and how many rows survive dropmissing once it is gone.
Where missingrows prices listwise deletion for the table as it stands, this prices the alternative — trading a variable for rows — and names the variable worth trading. The cells column (complete × columns left) sizes the surviving complete-case block; the step that maximizes it is flagged, since past that point each drop costs more in columns than it returns in rows.
Arguments
bar_width::Int: width in characters of the longest bar (default: 30).color::Symbol/missing_color::String: as inplotmissing.isna: predicate deciding what counts as absent (seeplotmissing).
Returns nothing; use missingdropstats for the same numbers as data.
MissingPatterns.missingdropstats — Methodmissingdropstats(tbl; isna=ismissing) -> Vector{<:NamedTuple}The listwise-deletion trade-off as a Tables.jl-compatible row table — the data behind missingdrop.
Complete-case analysis discards every row with a missing value, and a single sparse column can be responsible for most of that loss. This walks the greedy path of column drops — at each step removing the column that turns the most rows complete — and reports what complete-case analysis looks like after each one, so the cost of keeping a column is expressed in the rows it costs.
One row per step, starting from the untouched table, with fields:
ndropped::Int— columns dropped so far (0on the first row).dropped::Union{Nothing,Symbol}— the column dropped at this step;nothingon the first row, which is the table as given.ncols::Int— columns left.complete::Int— rows with no missing value among the columns left, i.e. whatdropmissingwould keep.pct::Float64—100 * complete / nrows.cells::Int—complete * ncols, the size of the complete-case block that survives. It rises while a drop buys more rows than it costs columns and falls afterwards, so its maximum is the natural stopping point.
The walk ends once every row is complete or one column is left. isna is the "counts as absent" predicate (see plotmissing).
Examples
using DataFrames
df = DataFrame(a=[1, 2, missing, 4], b=[1, missing, missing, 4], c=[1, 2, 3, 4])
DataFrame(missingdropstats(df))
# the drop that leaves the largest complete-case block
steps = missingdropstats(df)
best = steps[argmax([r.cells for r in steps])]See also missingrowstats for the distribution this optimizes against.
MissingPatterns.missinghtml — Methodmissinghtml(tbl; max_rows=200, max_cols=60, missing_color="#f3a9a9",
emphasis=:present, title="Missing data",
by=nothing, period=nothing, isna=ismissing,
order=:table) -> String
missinghtml(path::AbstractString, tbl; kwargs...) -> pathRender the missing-data heatmap as a standalone HTML fragment (dark-themed <div>, no external CSS/JS) — suitable for pasting into a blog post, notebook export, or report. Column headers are rotated for readability; every cell carries a tooltip with its row range and exact missing percentage. The same compression engine and color ramp as plotmissing are used, so the two outputs always agree — HTML just affords a much larger grid (defaults: 200×60 blocks).
by/period group rows by a column's values instead of by position, exactly as in plotmissing, so a grouped report reads the same in both media. isna and order likewise behave exactly as they do there.
The one-argument form returns the HTML String; the two-argument form writes it to path and returns the path.
MissingPatterns.missingpairstats — Methodmissingpairstats(tbl; isna=ismissing) -> Vector{<:NamedTuple}Pairwise co-occurrence of missingness as a Tables.jl-compatible row table — the data behind missingcooccurrence, without the rendering and without the max_cols display cap.
One row per unordered pair of distinct columns (ncols*(ncols-1)/2 rows, in upper-triangle order), with fields:
a::Symbol,b::Symbol— the two column names.phi::Float64— ϕ coefficient of the two missingness masks.jaccard::Float64— Jaccard index of the same masks.n11::Int— rows where bothaandbare missing.n1::Int,n2::Int— rows wherea(resp.b) is missing.nrows::Int— row count of the table.
Both coefficients fall out of the same n11/n1/n2 counts, so both are returned rather than selected by a method keyword: the schema stays fixed regardless of which one you look at.
Undefined coefficients come back as NaN, and the two differ on when that happens: phi is NaN whenever either column is entirely missing or entirely present (the 2x2 table is degenerate), while jaccard is NaN only when neither column has a single missing value — against a column with no missing values the union is still non-empty, so Jaccard is a well-defined 0.0.
Examples
pairs = missingpairstats(tbl)
# most co-missing pairs; `first` rather than `[1:5]`, which throws on a
# table with fewer than five pairs
first(sort(pairs; by = r -> -r.phi), 5)
filter(r -> r.n11 > 0 && r.jaccard > 0.5, pairs)See also missingstats, missingpatternstats.
MissingPatterns.missingpatterns — Methodmissingpatterns([io::IO=stdout], tbl; max_patterns=20, cell_chars=5,
char_missing='█', char_present='░', name_width=4,
color_cells=false, missing_color="#f3a9a9",
emphasis=:present, show_bar=true, min_pct=0.0,
isna=ismissing)Display the unique row-wise missingness patterns found in a Tables.jl-compatible tbl, sorted by descending frequency — i.e. which columns tend to be missing together.
This complements plotmissing, which shows where/how much is missing: missingpatterns shows which combinations of missing columns actually occur, the same diagnostic produced by R's mice::md.pattern(). Useful for reasoning about the missingness mechanism (e.g. if two columns are always missing together, that's rarely MCAR) and for choosing an imputation strategy. For a correlation-style view of the same question, see missingcooccurrence.
Arguments
io::IO: output stream (default:stdout).tbl: any Tables.jl-compatible table.max_patterns::Int: maximum number of patterns to display, most-frequent first (default: 20). Patterns beyond this are summarized in a trailing count.cell_chars::Int: number of repeated characters per cell (default: 5, max: 80).char_missing::Char: character for a missing column in a pattern (default:'█').char_present::Char: character for a present column in a pattern (default:'░').name_width::Int: max characters shown for column names before truncating with…(default: 4; set to 0 to show full names, bounded by cell width).color_cells::Bool: apply the same color ramp asplotmissing(default:false).missing_color::String: hex color of the ramp, as inplotmissing(default:"#f3a9a9").emphasis::Symbol::present(present columns colored, missing dark) or:missing(inverted), as inplotmissing(default::present).show_bar::Bool: append an UpSet-style horizontal frequency bar per pattern, scaled to the most common displayed pattern (default:true).min_pct::Float64: hide patterns matching fewer than this percentage of rows (default:0.0— show all). Hidden patterns are reported in the trailing summary line.isna: predicate deciding what counts as an absent value (default:ismissing), as inplotmissing.
Returns
nothing. The table is written toio.
MissingPatterns.missingpatternstats — Methodmissingpatternstats(tbl; isna=ismissing) -> Vector{<:NamedTuple}Unique row-wise missingness patterns as a Tables.jl-compatible row table — the data behind missingpatterns, without the rendering and without the max_patterns/min_pct display caps: every pattern is returned.
One row per distinct pattern, most frequent first (ties broken by first appearance, matching the displayed order), with fields:
pattern::NamedTuple— oneBoolper column, keyed by column name;truemeans missing in this pattern.nmissing::Int— how many columns are missing in the pattern.n::Int— rows matching it.pct::Float64—100 * n / nrows.
Examples
tbl = (age = [34, missing, 51], income = [missing, 4200, 5100])
ps = missingpatternstats(tbl)
ps[1].pattern.age # was `age` missing in the most frequent pattern?
filter(r -> r.nmissing == 0, ps) # the fully-complete pattern, if anypattern is a NamedTuple with one field per column, so its type depends on the table's width. On tables with many hundreds of columns this costs noticeable compile time on first call; the counts themselves stay cheap, since they come from the deduplicated pattern table.
See also missingstats, missingpairstats.
MissingPatterns.missingreport — Methodmissingreport(tbl; kwargs...) -> MissingReportWrap tbl in an object that renders itself as the missing-value heatmap in whichever medium displays it:
MIME"text/plain"(REPL, logs, files) → the terminal heatmap, identical toplotmissing.MIME"text/html"(Jupyter, Pluto, Documenter, any HTML-aware display) → the HTML heatmap, identical tomissinghtml.
So in a notebook missingreport(df) shows the colored HTML grid with per-cell tooltips, and the very same expression in a terminal shows the Unicode grid — with no if on the caller's side.
Keyword arguments are those of plotmissing and missinghtml; each is forwarded only to the renderer that accepts it, so per-medium defaults (e.g. a 200×60 HTML grid vs a 50×20 terminal grid) are preserved unless you override them explicitly. isna and order reach both renderers, so a report reads the same in either medium. An unknown keyword is an error at construction time, not at display time.
Examples
missingreport(df)
missingreport(df; emphasis=:missing, missing_color="#ff6600")
missingreport(df; by=:region) # grouped in both media
missingreport(df; layout=:compact, title="Cohort A")
# force one medium explicitly
show(stdout, MIME"text/html"(), missingreport(df))plotmissing and missinghtml are unchanged and remain the direct, single-medium entry points.
MissingPatterns.missingrows — Methodmissingrows([io::IO=stdout], tbl; sortby=:nmissing, bar_width=30,
color=:auto, missing_color="#f3a9a9", isna=ismissing)Distribution of how many values are missing per row: how many rows are complete, how many are missing exactly one value, two, and so on.
This is the view missingsummary (per column) and missingpatterns (per combination) leave out, and it is the one that answers "what would listwise deletion cost me?" — the 0 line is the complete-case count, everything below it is what dropmissing would discard.
Arguments
sortby::Symbol::nmissing(ascending missing-count, default) or:rows(descending row count — most common shape first).bar_width::Int: width in characters of the longest bar (default: 30).color::Symbol/missing_color::String: as inplotmissing; bars are tinted by severity (nmissing / ncols), so complete rows carry the "present" color and fully-missing rows the full missing color.isna: predicate deciding what counts as an absent value (default:ismissing), as inplotmissing.
Returns nothing; use missingrowstats for the same numbers as data.
MissingPatterns.missingrowstats — Methodmissingrowstats(tbl; isna=ismissing) -> Vector{<:NamedTuple}Row-completeness distribution as a Tables.jl-compatible row table — the data behind missingrows.
One row per observed missing-count (counts that occur zero times are omitted), ascending, with fields:
nmissing::Int— number of missing values in such a row (0= complete).nrows::Int— how many rows oftblhave exactly that many.pct::Float64—100 * nrows / total rows.
Examples
rs = missingrowstats(df)
only(r.nrows for r in rs if r.nmissing == 0) # complete-case count
sum(r.nrows for r in rs if r.nmissing > 0) # rows lost to listwise deletionSee also missingstats for the transposed (per-column) view.
MissingPatterns.missingstats — Methodmissingstats(tbl; isna=ismissing) -> Vector{<:NamedTuple}Per-column missing-data statistics as a Tables.jl-compatible row table — the data behind missingsummary, without the rendering.
One row per column of tbl, in table order, with fields:
column::Symbol— the column's name.eltype::Type— the column's element type,Missingincluded (missingsummarydisplaysBase.nonmissingtypeof this).nmissing::Int,npresent::Int— cell counts.nrows::Int— row count of the table (same on every row; carried so a single row is self-contained after filtering).pct::Float64—100 * nmissing / nrows, or0.0for an empty table.
Examples
using DataFrames
df = DataFrame(a=[1, missing, 3], b=["x", "y", "z"])
DataFrame(missingstats(df)) # straight into a DataFrame
filter(r -> r.pct > 20, missingstats(df))See also missingpatternstats, missingpairstats, missingrowstats.
MissingPatterns.missingsummary — Methodmissingsummary([io::IO=stdout], tbl; bins=20, sortby=:missing,
color=:auto, missing_color="#f3a9a9", isna=ismissing)Per-column missing-data overview: name, element type, missing count, %, and a sparkline showing where along the rows the missing values concentrate — the row axis is split into bins equal blocks and each block maps to a bar height proportional to its missing fraction. A block with even one missing value renders at least the smallest bar (same visibility guarantee as plotmissing); a block with none renders blank.
Arguments
bins::Int: sparkline resolution (default: 20).sortby::Symbol::missing(descending missing count, default),:name, or:none(table order).color::Symbol/missing_color::String: as inplotmissing; the sparkline colors bars by their missing fraction.isna: predicate deciding what counts as an absent value (default:ismissing), as inplotmissing.
MissingPatterns.plotmissing — Methodplotmissing([io::IO=stdout], tbl; cell_chars=5, char_missing='█', char_present='░',
name_width=4, color_cells=false, show_row_range=false,
max_rows=50, max_cols=20,
layout=:auto, target_lines=28, color=:auto,
missing_color="#f3a9a9", emphasis=:present,
by=nothing, period=nothing, isna=ismissing, order=:table)Display a text-based heatmap of missing value patterns in any Tables.jl-compatible source (DataFrame, CSV.File, NamedTuple of vectors, ...). When the data exceeds the display limits, multiple rows/columns are grouped into a single cell using a Unicode block-character gradient (classic layout) or an ANSI-colored half-block encoding (compact layout).
Layouts
:classic— the original layout: one grid row per line, 3-line header, 6-line summary. Best in a full terminal with room to scroll.:compact— fits the entire plot (grid + header + summary) in at mosttarget_lineslines, so IDE/Jupyter output cells never truncate it. With color available, each output line encodes two grid rows via'▀'(foreground = top row, background = bottom row), doubling vertical resolution; without color it falls back to the glyph gradient at one row per line. Any block containing even a single missing value is rendered in a shade distinct from "fully present", so fine holes survive compression.:auto(default) — uses:classicwhen it fits withintarget_lines,:compactotherwise.
Grouping
by::Union{Nothing,Symbol,String}: name of a column. When set, rows are grouped by the values of that column (not by position), so the vertical axis becomes honest categories/calendar time instead of arbitrary row ranges. Rows whosebyvalue ismissingform a trailing∅group. Row labels are always shown in this mode. If there are more groups than fit the budget, consecutive groups (in sorted order) are merged and labeled as ranges (e.g.2004-2005,A-C).period::Union{Symbol,Nothing}:nothing(default) groups by thebycolumn's exact value — categorical grouping, works for any sortable column (String,Symbol,Int, ...). Set to:year,:quarter,:month,:weekor:dayto instead group by that calendar period of aDate/DateTimebycolumn.:weekfollows ISO-8601, so its label carries the ISO week-year, which at a year boundary can differ from the calendar year (2024-12-30is2025-W01).
Arguments
io::IO: output stream (default:stdout).tbl: any Tables.jl-compatible table.cell_chars::Int: number of repeated characters per heatmap cell (default: 5, max: 80).char_missing::Char: character for fully-missing cells (default:'█').char_present::Char: character for fully-present cells (default:'░').name_width::Int: max characters shown for column names before truncating with…(default: 4; set to 0 to show full names, bounded by cell width).color_cells::Bool: apply the color ramp to classic-layout glyphs. The compact half-block layout always colors its cells (default:false).show_row_range::Bool: display row-range (or period) labels in a left-hand column (default:false; forcedtruewhenbyis set).max_rows::Int: maximum display rows before compression in the classic layout (default: 50). Ignored by:compact, which derives its own limit fromtarget_lines.max_cols::Int: maximum display columns before compression (default: 20).layout::Symbol::auto,:classic, or:compact(default::auto).target_lines::Int: total line budget for the compact layout, including borders, header and summary (default: 28 — safely under typical IDE output-cell limits of ~30 lines).color::Symbol::auto(respectio's:colorproperty / TTY detection),:always(force ANSI codes — use this in VS Code/Jupyter notebooks, whose output cells render ANSI but whosestdoutis not a TTY), or:never(plain text — use when redirecting to a file).missing_color::String: hex color ("#rrggbb") of the ramp (default:"#f3a9a9").emphasis::Symbol: which side of the data carries the ink (default::present). With:present, present data is painted inmissing_colorand missing data fades to dark gray — holes read as dark gaps in a colored field. With:missing, the ramp is inverted. In both modes, any block containing even one missing value renders in a shade visibly different from a fully-present block.isna: what counts as an absent value (default:ismissing). Microdata often codes absence as a sentinel rather than asmissing—9/99for "ignored",""for a blank field — and this lets those count as holes without rewriting the table. Either a predicate applied to every column:plotmissing(df; isna = x -> ismissing(x) || x == 9 || x == "")or, better, a
NamedTuple/AbstractDictof per-column predicates, withismissingassumed for any column left out:plotmissing(df; isna = (criterio = x -> ismissing(x) || x == 9, cid = x -> ismissing(x) || x == ""))Prefer the per-column form. A sentinel belongs to a variable, not to a table:
9means "ignored" in a coded field but is a perfectly good age, and a blanket predicate would punch holes in every column that happens to hold the value. Naming a column the table does not have is an error, not a silently ignored entry.In either form, test
ismissingfirst and let||short-circuit:missing == 9ismissing, notfalse, and would be an error in a boolean context. The predicate applies to every count the package makes, thebycolumn included, so a sentinel there forms the∅group.order::Symbol: column order (default::table).:tablekeeps the table's own order;:missingputs the emptiest columns first;:namesorts alphabetically;:clusterplaces columns that go missing together side by side, which is usually what makes a block structure visible at all — table order scatters them.:clusterseriates the ϕ matrix and so costs one extra pass over the data; columns with no missing values carry no pattern and are appended at the end. Reordering is purely a display concern and never changes a number. When columns are compressed, a reordered group is labeled by its endpoint names rather than by positional indices, which would otherwise refer to display slots.
Returns
nothing. The plot is written toio.
MissingPatterns.plotmissingdiff — Methodplotmissingdiff([io::IO=stdout], before, after; cell_chars=5, name_width=4,
max_cols=20, target_lines=28, color=:auto,
missing_color="#f3a9a9", filled_color="#a9f3c1",
isna=ismissing)Compare the missingness of two same-shaped tables — typically the same dataset before and after an imputation step, or two releases of a periodic microdata file. Rendered in the compact layout (fits target_lines):
- neutral dark gray — block unchanged;
- tint of
missing_color— block got more missing (introduced holes); - tint of
filled_color— block got less missing (holes resolved).
With color, half-blocks encode two row blocks per line (as in plotmissing); without color, cells fall back to + (more missing), - (fewer) and · (unchanged) glyphs. The summary line reports exact cell-level counts of resolved and introduced missing values, computed by a row-aligned pass (not from block averages).
Both tables must have identical dimensions and column names, in order.
isna: predicate deciding what counts as an absent value (default:ismissing), as inplotmissing.