Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

Modifications of the BIC for order selection in finite mixture models

Дата публикации: 20-07-2026 00:00:00

Finite mixture models are ubiquitous in modern statistical modeling, and a recurring practical issue is choosing the model order. In Keribin (, , 49–66, 2000), the Bayesian information criterion (BIC) was proved consistent in mixtures, but under strong regularity, including high moments and high-order derivatives of the component density. We introduce the $$\nu $$-BIC and $$\epsilon $$-BIC, which weight the BIC penalty by negligibly small logarithmic factors immaterial in practice. This minor modification yields consistency under substantially weaker conditions, without differentiability and with mild moment assumptions, and we also give a misspecification result: when the truth lies outside the candidate family, any vanishing-penalty IC eventually selects a Kullback–Leibler optimal order among candidates. Finally, we clarify two limitations of consistent IC-based selection in mixtures: there is no universally minimal BIC-scale penalty within our sufficient conditions, and order consistency can conflict with minimax optimality in Hellinger risk. We illustrate the theory for Gaussian mixtures, non-differentiable Laplace mixtures, heavy-tailed -mixtures, and mixtures of regression models.

Основное содержимое страницы с новостью.

1 Introduction

Finite mixtures are a natural and ubiquitous class of models for modeling heterogeneous populations consisting of distinct subpopulations. Multiple volumes have been written regarding the application, analysis, and implementation of such models, including the volumes of Titterington et al. (1985); Lindsay (1995); Peel and McLachlan (2000a); Fruhwirth-Schnatter et al. (2019); Ng et al. (2019); Chen (2023), and Yao and Xiang (2024), among many others. When conducting mixture modeling, a frequently required task is that of order selection. In this text, we will take an information criteria (IC) approach following the tradition of Akaike (1974) as espoused in the texts of Anderson and Burnham (2004); Claeskens and Hjort (2008), and Konishi and Kitagawa (2008). We characterise the task of order selection and the IC method as follows.

Let \(\left( \varOmega ,\mathcal{A},\textrm{P}\right) \) be a probability space with expectation operator \(\textrm{E}\) and suppose that \(X:\varOmega \rightarrow \mathbb {X}\) is a random map with density function \(f_{0}:\mathbb {X}\rightarrow \mathbb {R}_{>0}\) with respect to dominating measure \(\mathfrak {m}\), in the sense that \(\textrm{P}\left( X\in \mathbb {A}\right) =\int _{\mathbb {A}}f_{0}\textrm{d}\mathfrak {m}\), for each \(\mathbb {A}\in \mathcal{B}\left( \mathbb {X}\right) \) (the Borel \(\sigma \)-algebra of \(\mathbb {X}\)). Let \(\phi :\mathbb {X}\times \mathbb {T}\rightarrow \mathbb {R}_{>0}\) be a probability density function (PDF) dependent on parameter \(\theta =\left( \vartheta _{1},\dots ,\vartheta _{m}\right) \in \mathbb {T}\subset \mathbb {R}^{m}\) in the sense that \(\int _{\mathbb {X}}\phi \left( x;\theta \right) \mathfrak {m}\left( \textrm{d}x\right) =1\), for each \(\theta \), and define the k-component mixture of \(\phi \)-densities as

$$ \begin{aligned} \mathcal{M}_{k}^{\phi }=\Bigg \{\,\mathbb {X}\ni x\mapsto f_{k}(x;\psi _{k})&=\sum _{z=1}^{k}\pi _{z}\phi (x;\theta _{z}):\\&\theta _{z}\in \mathbb {T},\ \pi _{z}\in [0,1],\ \sum _{z=1}^{k}\pi _{z}=1,\ z\in [k] \,\Bigg \}, \end{aligned} $$

where \(\left[ k\right] =\left\{ 1,\dots ,k\right\} \) and \(\psi _{k}=\left( \pi _{1},\dots ,\pi _{k},\theta _{1},\dots ,\theta _{k}\right) \in \mathbb {S}_{k}\subset \mathbb {R}^{\left( 1+m\right) k}\) (the parameter space of densities in \(\mathcal{M}_{k}^{\phi }\)). Here, \(\phi \) is referred to as the component PDF.

We suppose that \(f_{0}\in \mathcal{M}_{k_{0}}^{\phi }\) but \(f_{0}\notin \mathcal{M}_{k_{0}-1}^{\phi }\) for some \(k_{0}\in \left[ \bar{k}\right] \), and some sufficiently large \(\bar{k}\in \mathbb {N}\). Here, the extra stipulation that \(f_{0}\notin \mathcal{M}_{k_{0}-1}^{\phi }\) is required because \(\left( \mathcal{M}_{k}^{\phi }\right) _{k\in \left[ \bar{k}\right] }\) is a nested sequence in the sense that \(\mathcal{M}_{k}^{\phi }\subset \mathcal{M}_{k+1}^{\phi }\), and thus if \(f_{0}\in \mathcal{M}_{k_{0}}^{\phi }\), it also holds that \(f_{0}\in \mathcal{M}_{k_{0}+1}^{\phi }\), therefore the condition alone does not uniquely define \(k_{0}\), which we refer to as the true order of \(f_{0}\). We may interpret \(k_{0}\) as the order of the smallest (or most parsimonious) model \(\mathcal{M}_{k}^{\phi }\) that contains \(f_{0}\). Let \(\ell _{k}\left( \psi _{k}\right) =-\textrm{E}\left\{ \log f_{k}\left( X;\psi _{k}\right) \right\} \) denote the average negative log-density of \(f_{k}\left( \cdot ;\psi _{k}\right) \in \mathcal{M}_{k}^{\phi }\), for each \(k\in \left[ \bar{k}\right] \). Then, by virtue of the divergence property of the Kullback–Leibler divergence (cf. e.g., Csiszár (1995)), it also holds that

$$ k_{0}=\min \underset{k\in \left[ \bar{k}\right] }{\arg \min }\left\{ \min _{\psi _{k}\in \mathbb {S}_{k}}\ell _{k}\left( \psi _{k}\right) \right\} , $$

assuming that the minimum is well-defined in each case, which is a useful characterisation that we will exploit in the sequel.

Suppose that we observe a data set of n independent and identically distributed (IID) replicates of X: \(\textbf{X}_{n}=\left( X_{1},\dots ,X_{n}\right) \subset \mathbb {X}\). Then, the problem of order selection can be characterised as the construction of an estimator \(\hat{k}_{n}=\hat{k}\left( \textbf{X}_{n}\right) \) of \(k_{0}\) with desirable properties. In particular, we want \(\hat{k}_{n}\) to be consistent in the sense that

$$ \lim _{n\rightarrow \infty }\textrm{P}\left( \hat{k}_{n}=k_{0}\right) =1. $$

If we let \(\ell _{k,n}\left( \psi _{k}\right) =-n^{-1}\sum _{i=1}^{n}\log f_{k}\left( X_{i};\psi _{k}\right) \) denote the empirical average negative log-likelihood of \(f_{k}\left( \cdot ;\psi _{k}\right) \), the natural estimator of \(\ell _{k}\left( \psi _{k}\right) \), then the method of IC suggests that we construct the estimator of \(k_{0}\) in the form

$$ \hat{k}_{n}=\min \underset{k\in [\bar{k}]}{\arg \min }\left\{ \min _{\psi _{k}\in \mathbb {S}_{k}}\ell _{k,n}(\psi _{k})+\textrm{pen}_{k,n}\right\} , $$

for some appropriately chosen penalties \(\textrm{pen}_{k,n}\). Here, we may recognise \(\min _{\psi _{k}\in \mathbb {S}_{k}}\ell _{k,n}\left( \psi _{k}\right) \) as the negative maximum likelihood for the mixture model of order k.

We can easily identify this approach with some traditional methods, such as the Akaike information criterion (AIC; Akaike (1974)) and the Bayesian information criterion (BIC; Schwarz (1978)), by taking \(\textrm{pen}_{k,n}^{\textrm{AIC}}=d_{k}n^{-1}\) and \(\textrm{pen}_{k,n}^{\textrm{BIC}}=d_{k}\left( 2n\right) ^{-1}\log n\), respectively, where \(d_{k}\) denotes the parameter count used in the criterion. Let \(\textrm{Ln}\left( n\right) =\log \left( e\vee n\right) \) denote the logarithm function truncated to 1, from below, where \(a\vee b=\max \left\{ a,b\right\} \), and denote its \(\nu \)-fold composition by \(\textrm{Ln}^{\circ \nu }\left( n\right) \) (e.g., \(\textrm{Ln}^{\circ 3}\left( n\right) =\textrm{Ln}\circ \textrm{Ln}\circ \textrm{Ln}\left( n\right) \)). In this work, we introduce the \(\nu \)-BIC and \(\epsilon \)-BIC families of penalties of the forms \(\textrm{pen}_{k,n}^{\nu }=\alpha \left( k\right) n^{-1}\textrm{Ln}^{\circ \nu }\left( n\right) \log n\) and \(\textrm{pen}_{k,n}^{\epsilon }=\alpha \left( k\right) n^{-1}\left( \log n\right) ^{1+\epsilon }\), for any choice of \(\epsilon >0\) and \(\nu \in \mathbb {N}\), and for \(\alpha :\mathbb {N}\rightarrow \mathbb {R}_{>0}\) strictly increasing. A BIC-style calibration is obtained by taking \(\alpha \left( k\right) =d_{k}/2\). We shall demonstrate that these modest modifications of the BIC are consistent under weaker assumptions than those currently available for BIC consistency in mixtures, as we shall elaborate upon in the sequel. We also provide a complementary misspecification result: when \(f_{0}\) need not belong to any \(\mathcal{M}_{k}^{\phi }\), any IC satisfying a mild vanishing-penalty condition will eventually select an order whose best-fitting mixture model is Kullback–Leibler optimal among the candidate orders.

For regular models, order selection consistency of the BIC was proveds in the works of Nishii (1988); Vuong (1989), and Sin and White (1996). Unfortunately, the mixture model setting is not regular and thus alternative proofs of consistency are required. A prime and often cited example of such a result is that of Keribin (2000) who establish the consistency of the BIC for mixture models under strong assumptions. Namely, among other things, the proof requires that the component density have up to five derivatives and that all of these higher derivatives have finite third moments. A remarkable result of Gassiat and Van Handel (2012) (see also Gassiat 2018, Sec. 4.3) shows that the BIC is order consistent for choosing between models when \(\bar{k}=\infty \), albeit at the price of much stronger assumptions. In particular, even when restricted to the class of location mixtures, the proof requires that \(\phi \) has up to three derivatives and that all derivatives up to second order have finite third moments, while third order derivatives have second moments.

Beyond the BIC, Drton and Plummer (2017) proposed the so-called singular BIC (sBIC) as an approach that operationalises the expansion \(n\ell _{k,n}\left( \psi _{k}\right) +\lambda _{k}\left( f_{0}\right) \log n+\left\{ m_{k}\left( f_{0}\right) -1\right\} \log \log n+O_{\textrm{P}}\left( 1\right) \) of the so-called Bayes free energy of Watanabe (2013), related to the widely-applicable BIC (WBIC), where the coefficients \(\left( \lambda _{k}\left( f_{0}\right) ,m_{k}\left( f_{0}\right) \right) _{k\in \left[ \bar{k}\right] }\) are dependent on the true and unknown model \(f_{0}\). As noted in Chen (2023, Sec. 16.3), the specification of the sBIC arises as a solution to a nonlinear system of stochastic equations making it challenging to faithfully implement in practice. In Baudry (2015), the log-likelihood is replaced by the so-called complete data log-likelihood and a consistency result is obtained with the usual BIC-form penalty. However, the proof makes the requirement that the limiting objective functionals of all order \(k\in \left[ \bar{k}\right] \) have positive definite Hessian, which is not possible in the context of mixture models, as typically if \(k_{0}<k\), then the minimisers of any risk over \(\mathcal{M}_{k}^{\phi }\) will have connected components. In the context of test-based order selection, this phenomenon is discussed in Chen (2023, Ch. 9).

Sacrificing the requirement that \(\textrm{pen}_{k,n}=\tilde{O}\left( n^{-1}\right) \) (where we use the soft-Oh notation \(\tilde{O}\) to denote the Landau order up to logarithmic factors; cf. Cormen et al. (2002), Sec. 3-5), Nguyen (2024) propose to take penalties such that \(\lim _{n\rightarrow \infty }\sqrt{n}\textrm{pen}_{k,n}=\infty \), e.g. \(\textrm{pen}_{k,n}=\tilde{O}\left( n^{-1/2}\right) \), within their PanIC framework for model selection in general risk minimisation problems (see also Westerhout et al. (2024)). Using these larger penalties, it is possible to prove model selection consistency under very mild conditions, with data potentially dependent, and where the objective functions (including but not necessarily log-likelihoods) may be non-differentiable or even discontinuous and unmeasurable. Although these results are very general, it is clear that for sufficiently well-behaved models, \(\tilde{O}\left( n^{-1/2}\right) \) penalties are too large and as a consequence, the corresponding IC \(\hat{k}_{n}\) may be inefficient and thus larger than necessary values of n are required for consistency to manifest, in practice.

In this work, we will prove that the slightly larger penalties \(\textrm{pen}_{k,n}^{\nu }\) and \(\textrm{pen}_{k,n}^{\epsilon }\) allow for a substantial relaxation of the conditions required for order consistency. In particular, we will show that consistency can be proved when the parameter space \(\mathbb {S}_{k}\) is compact, the log-density class is a Glivenko–Cantelli (GC) class, and the component density functions are Lipschitz over \(\mathbb {T}\) with an \(\mathcal {L}_{1}(\mathfrak m)\) envelope. These conditions are far less onerous than those of Keribin (2000) and allow for application in cases where the component densities \(\phi \) need not even be differentiable. In fact, penalties \(\textrm{pen}_{k,n}^{\epsilon }\) were considered within the framework of Keribin (2000) but the larger penalisation was not exploited to weaken the regularity conditions.

In fact, for \(\nu \) chosen sufficiently large or \(\epsilon \) chosen sufficiently small, the decision made by using the \(\nu \)-BIC or \(\epsilon \)-BIC versus that by the BIC will be the same, in practical settings. For instance, when \(\nu =3\), \(\textrm{Ln}^{\circ \nu }\left( n\right) \le 1\) for all \(n\le \exp ^{\circ 3}\left( 1\right) \approx 3.8\times 10^{6}\) and \(\textrm{Ln}^{\circ \nu }\left( n\right) \le 1.1\) for all \(n\le \exp ^{\circ 3}\left( 1.1\right) \approx 5.7\times 10^{8}\). Similarly, taking \(\epsilon =0.02\), \(\left( \log n\right) ^{0.02}\le 1.1\) for all \(n\le \exp \left( 117.3909\right) \approx 9.6\times 10^{50}\). Thus, one key interpretation of our results is that they provide theoretical support for penalties that are numerically very close to the BIC outside the regularity framework of Keribin (2000). In finite sample studies, the use of the BIC may therefore be viewed as numerically close to the use of \(\nu \)-BIC or \(\epsilon \)-BIC penalties, up to a negligible multiplicative factor. We also use our framework to clarify two limitations of consistent IC-based order selection in mixtures. First, within the general sufficient conditions used to establish consistency, there is no universally minimal BIC-scale penalty, so consistency alone does not single out an optimal IC. Second, we establish a tension between order consistency and minimax optimality in mixtures in the style of Yang (2005). In particular, we show that in a simple Gaussian mixture family, an AIC-like criterion achieves the parametric minimax Hellinger risk, whereas any order-consistent criterion (including the BIC and our proposals) must be minimax-inefficient by an unbounded factor. To the best of our knowledge, this mixture model AIC/BIC tension in mixture models has not previously been stated explicitly.

Indeed, the method of IC is not the only approach for order selection and estimation under model uncertainty in mixture models. For example, model selection using hypothesis testing has been considered in McLachlan (1987); Polymenis and Titterington (1998); Wasserman et al. (2020); Nguyen et al. (2023), and the many approaches covered in Chen (2023). Shrinkage estimator approaches have been proposed in Chen and Khalili (2009); Huang et al. (2017, 2022), and Budanova (2025). And bounds of risk functions for estimation under model uncertainty have been considered by Li and Barron (1999); Rakhlin et al. (2005); Klemelä (2007); Maugis and Michel (2011); Manole and Ho (2022), and Chiu Chong et al. (2024), among many others. We do not comment or compare these disparate approaches to model selection, and direct the reader to dedicated texts, such as Peel and McLachlan (2000a, Ch. 6), McLachlan and Rathnayake (2014), Fruhwirth-Schnatter et al. (2019, Ch. 7), and Chen (2023, Ch. 16).

To conclude, we note that beyond mixtures of density functions, one can consider IC for conditional density models, as well. Results in this setting include the works of Naik et al. (2007); Hafidi and Mkhadri (2010); Depraetere and Vandebroek (2014), and Hui et al. (2015). It is possible to adapt our approach to provide inference in this setting, and we will consider such situations as an example application of our theory.

The remainder of the manuscript is organised as follows. Section 2 introduces the notational conventions used throughout the paper, presents the technical preliminaries, and states our main consistency results, including a misspecification result in which order selection targets Kullback–Leibler optimal approximating models. Section 3 provides illustrative examples that demonstrate applications of our theory. Section 4 then discusses limitations of consistent model selection in mixture models, including penalty flexibility and the aforementioned AIC/BIC tension. Concluding remarks and discussion are then made in Sect. 5. Complete proofs of our technical results and primary theorems are provided in the supplementary material. Auxiliary assumptions used for comparison are provided in the Appendix.

2 Main results
2.1 Notation and conventions

We retain the setup and notation from Sect. 1 and follow the usual notational conventions for empirical processes (see, e.g., van de Geer (2000)). In particular, \(\mathcal{M}_{k}^{\phi }\) denotes the class of \(k\)-component mixture densities, \(f_{k}(\cdot ;\psi _{k})\) denotes a density in that class, \(\mathbb {S}_{k}\) denotes its parameter space, \(\ell _{k}\) and \(\ell _{k,n}\) denote the population and empirical average negative log-likelihoods, \(\textrm{pen}_{k,n}\) denotes the order-\(k\) penalty, and \(\hat{k}_{n}\) denotes the resulting IC-selected order. Let \(P\) denote the true probability measure of \(X\), and use empirical-process notation \(Pf=\textrm{E}f(X)\) for an integrable function \(f\). Let \(P_n\) denote the empirical measure, so that \(P_nf=n^{-1}\sum _{i=1}^n f(X_i)\). For a functional class \(\mathcal{F}\), \(\ell ^\infty (\mathcal{F})\) is the space of bounded functionals on \(\mathcal{F}\) equipped with the uniform norm \(\Vert h\Vert _{\mathcal{F}}=\sup _{f\in \mathcal{F}}|h(f)|\). When the class is clear from context, \(P_n-P\) is viewed as an element of \(\ell ^\infty (\mathcal{F})\). A class \(\mathcal{F}\) is GC if it satisfies the uniform strong law of large numbers \(\Vert P_n-P\Vert _\mathcal {F}\xrightarrow [n\rightarrow \infty ]{\mathrm {a.s.}}0\). We write \(a_n\asymp b_n\) if there exist constants \(0<c\le C<\infty \) and \(n_0\in \mathbb {N}\) such that \(cb_n\le a_n\le Cb_n\) for all \(n\ge n_0\), and write \(g(n)\ll h(n)\) if \(g(n)/h(n)\rightarrow 0\).

For \(p\ge 1\), \(\mathcal{L}_p(P)=\mathcal{L}_p(\mathbb {X},\mathcal{B}(\mathbb {X}),P)\) denotes the usual Lebesgue space, with its associated norm. For a vector \(\psi \in \mathbb {R}^q\), \(\Vert \psi \Vert \) denotes the Euclidean norm and \(\Vert \psi \Vert _p\) denotes the \(\ell _p\) norm. For densities \(f\) and \(g\) with respect to \(\mathfrak {m}\), the Hellinger distance is

$$ \mathfrak {h}(f,g)=\left\{ \frac{1}{2}\int _{\mathbb {X}}(f^{1/2}-g^{1/2})^2\,\textrm{d}\mathfrak {m}\right\} ^{1/2}. $$

For probability measures \(Q\) and \(R\) on \((\mathbb {X},\mathcal{B}(\mathbb {X}))\), the Kullback–Leibler divergence is

$$ \mathfrak {K}(Q\Vert R)={\left\{ \begin{array}{ll} \displaystyle \int _{\mathbb {X}}\log \left( \frac{\textrm{d}Q}{\textrm{d}R}\right) \,\textrm{d}Q & \text {if }Q\ll R,\\ \infty & \text {otherwise.} \end{array}\right. } $$

If \(Q\) and \(R\) have densities \(f\) and \(g\) with respect to \(\mathfrak {m}\), we also write

$$ \mathfrak {K}(f\Vert g)=\int _{\mathbb {X}} f(x)\log \{f(x)/g(x)\}\mathfrak {m}(\textrm{d}x). $$

We write

$$ L_k=\min _{\psi _k\in \mathbb {S}_k}\ell _k(\psi _k),\qquad \hat{L}_{k,n}=\min _{\psi _k\in \mathbb {S}_k}\ell _{k,n}(\psi _k), $$

for the corresponding minimized risks. Thus \(L_k\) and \(\hat{L}_{k,n}\) denote the minimized population and empirical risks, respectively, as distinct from the parameter-indexed risks \(\ell _k\) and \(\ell _{k,n}\). When likelihood comparisons are written in maximisation form over density classes \(\mathcal {F}_k\), we use, for any probability measure \(Q\) for which the displayed quantity is well defined,

$$ \varLambda _{k,n}=\sup _{f\in \mathcal {F}_k}P_n\log f=-\inf _{f\in \mathcal {F}_k}\ell _n(f), \qquad \varLambda _k(Q)=\sup _{f\in \mathcal {F}_k}Q\log f. $$

When the relevant law is clear, we write \(\varLambda _k\) in place of \(\varLambda _k(Q)\). Consequently, when \(\mathcal {F}_k=\mathcal{M}_k^\phi \), one has \(\hat{L}_{k,n}=-\varLambda _{k,n}\) and \(L_k=-\varLambda _k(P)\).

For a norm \(\Vert \cdot \Vert \), a collection \(( [f_j^{\mathrm L},f_j^{\mathrm U}] )_{j\in [N]}\) is a \(\delta \)-bracketing of \(\mathcal{F}\) if, for each \(f\in \mathcal{F}\), there is a \(j\) such that \(f_j^{\mathrm L}\le f\le f_j^{\mathrm U}\) and \(\Vert f_j^{\mathrm U}-f_j^{\mathrm L}\Vert \le \delta \), for every \(j\in [N]\). The bracketing number \(N_{\left[ \right] }(\delta ,\mathcal{F},\Vert \cdot \Vert )\) is the smallest such \(N\), with value \(\infty \) if no finite bracketing exists, and \(H_{\left[ \right] }(\delta ,\mathcal{F},\Vert \cdot \Vert )=\log N_{\left[ \right] }(\delta ,\mathcal{F},\Vert \cdot \Vert )\) is the corresponding bracketing entropy. When \(\mathcal{F}\subset \mathcal{L}_p(P)\) and \(\Vert \cdot \Vert \) is the \(\mathcal{L}_p(P)\) norm, we write \(N_{\left[ \right] }(\delta ,\mathcal{F},\mathcal{L}_p(P))\) for short. For a metric space \((\mathcal{G},d)\), a collection \((g_j)_{j\in [N]}\) is a \(\delta \)-covering of \(\mathcal{G}\) if each \(g\in \mathcal{G}\) is within distance \(\delta \) of some \(g_j\); the smallest such \(N\) is denoted \(N(\delta ,\mathcal{G},d)\).

2.2 Technical preliminaries

With this notation, the following results are useful for proving our technical outcomes.

Lemma 1

If \(\textbf{X}_{n}\) is IID and \(\mathcal{F}\subset \mathcal{L}_{1}\left( P\right) \) is a class of measurable functions admitting an envelope \(F\in \mathcal{L}_{1}\left( P\right) \) and \(N_{\left[ \right] }\left( \delta ,\mathcal{F},\mathcal{L}_{1}\left( P\right) \right) <\infty \), for every \(\delta >0\), then \(\mathcal{F}\) is a GC class. In particular, if

$$\begin{aligned} \mathcal{F}=\left\{ f_{g}:\mathbb {X}\ni x\mapsto f_{g}\left( x\right) ,g\in \mathcal{G}\right\} , \end{aligned}$$

(1)

\(\left( \mathcal{G},d\right) \) is compact and \(g\mapsto f_{g}\left( x\right) \) is continuous for P-almost all \(x\in \mathbb {X}\), then \(N_{\left[ \right] }\left( \delta ,\mathcal{F},\mathcal{L}_{1}\left( P\right) \right) <\infty \), for every \(\delta >0\), if there exists an envelope function \(F\in \mathcal{L}_{1}\left( P\right) \), such that \(\sup _{g\in \mathcal{G}}\left| f_{g}\right| \le F\).

Proof

See van de Geer (2000, Thm. 2.4.1 and Lem. 3.10). \(\square \)

Remark 1

As an alternative to the bracketing entropy condition, above, one can also verify that \(\mathcal{F}\) is a GC class via random covering entropy methods (see, e.g., van de Geer (2000), Sec. 3.6). We do not use such conditions in this text since bracketing entropy bounds pair better with the rest of our theoretical approach as shall be seen, in the sequel.

Remark 2

Throughout this text, we shall assume that all random quantities of interest are measurable. Indeed, all of the classes \(\mathcal{F}\) considered in this work can be understood as classes of Carathéodory integrands that are indexed by Euclidean spaces and thus are measurable, with measurable maxima/minima, sums, compositions, and other required manipulations (cf. Rockafellar and Wets (2009), Ch. 14).

Lemma 2

If \(\mathcal{F}\) is parameterised as per (1), such that there exists an envelope function \(F:\mathbb {X}\rightarrow \mathbb {R}_{>0}\) satisfying

$$ \left| f_{g}\left( x\right) -f_{h}\left( x\right) \right| \le d\left( g,h\right) F\left( x\right) , $$

then \(N_{\left[ \right] }\left( 2\delta \left\| F\right\| ,\mathcal{F},\left\| \cdot \right\| \right) \le N\left( \delta ,\mathcal{G},d\right) \). Furthermore, if \(\mathcal{G}=\mathbb {S}\subset \mathbb {R}^{q}\) is compact and d is the Euclidean distance, then there exists a constant \(K>0\) such that

$$ N\left( \delta ,\mathbb {S},d\right) \le K\left( \frac{\textrm{diam}\left( \mathbb {S}\right) }{\delta }\right) ^{q}, $$

for every \(0<\delta <\textrm{diam}\left( \mathbb {S}\right) \).

Proof

See van der Vaart and Wellner (van der Vaart and Wellner (2023), Thm. 2.7.17) and van der Vaart (2000, Example 19.7). \(\square \)

Next, suppose that \(\mathcal{F}\) is a subset of probability density functions on \(\mathbb {X}\), in the sense that for each \(f\in \mathcal{F}\), \(f\left( x\right) >0\) for each \(x\in \mathbb {X}\) and \(\int _{\mathbb {X}}f\textrm{d}\mathfrak {m}=1\). For each \(n\in \mathbb {N}\), we let \(\hat{f}_{n}:\varOmega \rightarrow \mathcal{F}\) be a maximum likelihood estimator (MLE), defined by

$$\begin{aligned} \hat{f}_{n}\in \underset{f\in \mathcal{F}}{\arg \max }P_{n}\log f, \end{aligned}$$

(2)

where the existence of MLEs is guaranteed in our application. Using the Hellinger distance defined in Sect. 2.1, write the subset of \(\mathcal{F}\) that lies within the \(\delta \)-Hellinger ball of \(f_{0}\), for \(0<\delta \le 1\), as \(\mathcal{F}\left( \delta \right) =\left\{ f\in \mathcal{F}:\mathfrak {h}\left( f,f_{0}\right) \le \delta \right\} \). Further define the following entropy integral:

$$\begin{aligned} J\left( \delta \right) =\delta \vee \int _{0}^{\delta }\sqrt{H_{\left[ \right] }\left( u^{2},\mathcal{F}\left( \delta \right) ,\mathcal{L}_{1}\left( \mathfrak {m}\right) \right) }\textrm{d}u. \end{aligned}$$

(3)

The next result, our main technical tool, is adapted from van de Geer (2000, Cor. 7.5) (see the supplementary material).

Proposition 1

If \(f_{0}\in \mathcal{F}\), \(\textbf{X}_{n}\) is IID and there is some \(\varPsi :\mathbb {R}_{>0}\rightarrow \mathbb {R}_{>0}\) such that \(\varPsi \left( \delta \right) \ge J\left( \delta \right) \), where \(\varPsi \left( \delta \right) /\delta ^{2}\) is a non-increasing function of \(\delta \), then, for some constant \(c_{1}>0\),

$$ \textrm{P}\left( \int _{\mathbb {X}}\log \frac{\hat{f}_{n}}{f_{0}}\textrm{d}P_{n}\ge \delta ^{2}\right) \le c_{1}\exp \left\{ -\frac{n\delta ^{2}}{c_{1}^{2}}\right\} , $$

for every \(\delta \ge \delta _{n}\), such that \(\delta _{n}\) satisfies \(\sqrt{n}\delta _{n}^{2}\ge c\varPsi \left( \delta _{n}\right) \), for a universal constant \(c>0\).

2.3 Assumptions

We state the following assumptions, repeating some previously stated items, for convenience.

A1:

\(X\in \mathbb {X}\) has PDF \(f_{0}\in \mathcal{M}_{k_{0}}^{\phi }\setminus \mathcal{M}_{k_{0}-1}^{\phi }\), for some \(k_{0}\in \left[ \bar{k}\right] \) (with the convention \(\mathcal{M}_{0}^{\phi }=\varnothing \)), and \(\textbf{X}_{n}=\left( X_{1},\dots ,X_{n}\right) \) is a sample of IID replicates of X.

A2:

\(\phi \) is strictly positive; i.e., \(\phi \left( x;\theta \right) >0\), for each \(x\in \mathbb {X}\) and \(\theta \in \mathbb {T}\subset \mathbb {R}^{m}\), where \(\mathbb {T}\) is compact.

A3:

\(\phi \) is a Carathéodory integrand in the sense that \(\phi \left( x;\cdot \right) :\mathbb {T}\rightarrow \mathbb {R}\) is continuous for each fixed \(x\in \mathbb {X}\), and \(\phi \left( \cdot ;\theta \right) :\mathbb {X}\rightarrow \mathbb {R}\) is measurable, for each fixed \(\theta \in \mathbb {T}\).

A4:

There exists a \(G_{1}\in \mathcal{L}_{1}\left( P\right) \) such that, for every \(x\in \mathbb {X}\),

$$ \max _{\theta \in \mathbb {T}}\left| \log \phi \left( x;\theta \right) \right| <G_{1}\left( x\right) . $$

A5:

There exist functions \(L_{\phi }:\mathbb {X}\rightarrow \mathbb {R}_{\ge 0}\) and \(G_{2}\in \mathcal{L}_{1}\left( \mathfrak {m}\right) \) such that, for every \(\theta _{1},\theta _{2}\in \mathbb {T}\) and \(x\in \mathbb {X}\),

$$\begin{aligned} \left| \phi \left( x;\theta _{1}\right) -\phi \left( x;\theta _{2}\right) \right| \le L_{\phi }\left( x\right) \left\| \theta _{1}-\theta _{2}\right\| _{1}, \end{aligned}$$

(4)

where \(\max _{\theta \in \mathbb {T}}\phi \left( x;\theta \right) +L_{\phi }\left( x\right) \le G_{2}\left( x\right) \).

When \(\mathbb {T}\) is convex and \(\phi \left( x,\cdot \right) :\mathbb {T}\rightarrow \mathbb {R}\) is continuously differentiable, for each \(x\in \mathbb {X}\), we can replace A5 with the following condition.

A5*:

The set \(\mathbb {T}\) is convex, and there exists a function \(G_{2}\in \mathcal{L}_{1}\left( \mathfrak {m}\right) \) such that, for each \(x\in \mathbb {X}\),

$$ \max _{\theta \in \mathbb {T}}\phi \left( x;\theta \right) +\max _{j\in \left[ m\right] }\max _{\theta \in \mathbb {T}}\left| \frac{\partial \phi \left( x;\theta \right) }{\partial \vartheta _{j}}\right| \le G_{2}\left( x\right) . $$

We also impose the following penalty conditions, under which penalties are non-negative throughout.

B1:

For each \(k\in \left[ \bar{k}\right] \), \(\textrm{pen}_{k,n}\ge 0\) for all n and \(\lim _{n\rightarrow \infty }\textrm{pen}_{k,n}=0\).

B2:

For each \(1\le k<l\le \bar{k}\),

$$ \lim _{n\rightarrow \infty }\frac{n}{\log n}\left\{ \textrm{pen}_{l,n}-\textrm{pen}_{k,n}\right\} =\infty . $$

Let us provide some reasoning for the assumptions. A1 identifies that the true PDF \(f_{0}\) is within the class of mixture models \(\mathcal{M}_{k}^{\phi }\), where \(k\le \bar{k}\). We also have the availability of an IID sample of n replicates \(\textbf{X}_{n}\) from the true probability model, required for application of Lemma 1 and Proposition 1. A2 provides the necessary compactness of the parameter space for bounding the metric entropy using Lemma 2 and ensures that the log-densities \(\log \phi \left( x;\theta \right) \ne \infty \), for each pair of x and \(\theta \). A3 implies that the density expressions, log-density expressions, and their minima/maxima and averages remain measurable. A4 provides envelope functions so that we can apply Lemma 1 to prove the convergence of the average log-likelihoods to their limits, for each k. And A5 provides the required envelope functions for bounding the bracketing entropy using Lemma 2. Note that when \(\phi \) is continuously differentiable in \(\theta \), for each fixed x, and \(\mathbb {T}\) is convex, A5* implies A5 by the mean value theorem. B1 imposes non-negative vanishing penalties and guarantees that the penalty does not interfere with comparisons between models \(\mathcal{M}_{k}^{\phi }\) and \(\mathcal{M}_{k_{0}}^{\phi }\), with respect to their maximum average log-likelihoods, when \(k<k_{0}\). B2 then provides sufficient penalisation to distinguish between \(\mathcal{M}_{k}^{\phi }\) and \(\mathcal{M}_{k_{0}}^{\phi }\), with respect to their model complexities, when \(k>k_{0}\).

2.4 Consistency results

The following pair of lemmas, proved in the supplementary material, constitute the main technical contribution of the text and will imply the main result.

Lemma 3

Under A1–A4, for each \(k\in \left[ \bar{k}\right] \),

$$ \min _{\psi _k\in \mathbb {S}_k}\ell _{k,n}(\psi _k) \xrightarrow [n\rightarrow \infty ]{\mathrm {a.s.}} \min _{\psi _k\in \mathbb {S}_k}\ell _k(\psi _k). $$

Lemma 4

Under A1–A3 and A5, for each \(k,l\ge k_{0}\),

$$ \frac{n}{\log n}\left\{ \min _{\psi _{k}\in \mathbb {S}_{k}}\ell _{k,n}\left( \psi _{k}\right) -\min _{\psi _{l}\in \mathbb {S}_{l}}\ell _{l,n}\left( \psi _{l}\right) \right\} =O_{\textrm{P}}\left( 1\right) . $$

As a consequence, we have the following generic consistency result and our main result as a corollary.

Theorem 1

Under A1–A5, B1, and B2, \(\lim _{n\rightarrow \infty }\textrm{P}\left( \hat{k}_{n}=k_{0}\right) =1\).

Proof

Write \(\hat{L}_{k,n}=\min _{\psi _{k}\in \mathbb {S}_{k}}\ell _{k,n}(\psi _{k})\) and \(L_{k}=\min _{\psi _{k}\in \mathbb {S}_{k}}\ell _{k}(\psi _{k})\). Since \(f_{0}\in \mathcal{M}_{k_{0}}^{\phi }\setminus \mathcal{M}_{k_{0}-1}^{\phi }\) and the model sequence is nested, \(L_{k}>L_{k_{0}}\) for every \(k<k_{0}\). Indeed, \(L_{k}-L_{k_{0}}=\inf _{f\in \mathcal{M}_{k}^{\phi }}\mathfrak {K}(P\Vert P_{f})\), where \(P_{f}\) denotes the measure with density f. If the infimum were zero, Pinsker’s inequality and compactness of the parameter space, together with the \(\mathcal{L}_{1}(\mathfrak {m})\)-continuity supplied by A5, would imply that \(f_{0}\in \mathcal{M}_{k}^{\phi }\), contradicting A1. Because there are finitely many such k, there exists \(\eta >0\) such that \(L_{k}-L_{k_{0}}>2\eta \) for all \(k<k_{0}\). Lemma 3 and B1 imply that, with probability tending to one, \(\hat{L}_{k,n}+\textrm{pen}_{k,n}>\hat{L}_{k_{0},n}+\textrm{pen}_{k_{0},n}\) for every \(k<k_{0}\).

For \(k>k_{0}\), Lemma 4 gives

$$ \frac{n}{\log n}\left( \hat{L}_{k,n}-\hat{L}_{k_{0},n}\right) =O_{\textrm{P}}(1), $$

whereas B2 gives

$$ \frac{n}{\log n}\left( \textrm{pen}_{k,n}-\textrm{pen}_{k_{0},n}\right) \rightarrow \infty . $$

Hence, with probability tending to one, \(\hat{L}_{k,n}+\textrm{pen}_{k,n}>\hat{L}_{k_{0},n}+\textrm{pen}_{k_{0},n}\) for every \(k>k_{0}\). Since \(\bar{k}\) is finite, combining the underfitting and overfitting comparisons yields \(\textrm{P}(\hat{k}_{n}=k_{0})\rightarrow 1\). \(\square \)

Corollary 1

Under A1–A5, for any strictly increasing \(\alpha :\mathbb {N}\rightarrow \mathbb {R}_{>0}\), \(\nu \in \mathbb {N}\) and \(\epsilon >0\), the \(\nu \)-BIC

$$ \hat{k}_{n}^{\nu }=\min \arg \min _{k\in \left[ \bar{k}\right] }\left\{ \min _{\psi _{k}\in \mathbb {S}_{k}}\ell _{k,n}\left( \psi _{k}\right) +\alpha \left( k\right) \frac{\textrm{Ln}^{\circ \nu }\left( n\right) \log n}{n}\right\} $$

and \(\epsilon \)-BIC

$$ \hat{k}_{n}^{\epsilon }=\min \arg \min _{k\in \left[ \bar{k}\right] }\left\{ \min _{\psi _{k}\in \mathbb {S}_{k}}\ell _{k,n}\left( \psi _{k}\right) +\alpha \left( k\right) \frac{\log ^{1+\epsilon }n}{n}\right\} $$

are consistent estimators of \(k_{0}\).

Proof

For every fixed \(\nu \in \mathbb {N}\) and \(\epsilon >0\), both \(\textrm{Ln}^{\circ \nu }\left( n\right) \log n=o(n)\) and \(\log ^{1+\epsilon }n=o(n)\). Thus \(\textrm{pen}_{k,n}^{\nu }\) and \(\textrm{pen}_{k,n}^{\epsilon }\) satisfy B1. For \(k<l\), observe that

$$ \frac{n}{\log n}\left\{ \textrm{pen}_{l,n}^{\nu }-\textrm{pen}_{k,n}^{\nu }\right\} =\left\{ \alpha \left( l\right) -\alpha \left( k\right) \right\} \textrm{Ln}^{\circ \nu }\left( n\right) \underset{n\rightarrow \infty }{\longrightarrow }\infty , $$

and

$$ \frac{n}{\log n}\left\{ \textrm{pen}_{l,n}^{\epsilon }-\textrm{pen}_{k,n}^{\epsilon }\right\} =\left\{ \alpha \left( l\right) -\alpha \left( k\right) \right\} \log ^{\epsilon }n\underset{n\rightarrow \infty }{\longrightarrow }\infty , $$

as required, since \(\log \) and positive powers are increasing, and since \(\alpha \) is strictly increasing. \(\square \)

To illustrate the advantage of Theorem 1 over the main result of Keribin (2000) (i.e., Theorem 2.1), it is helpful to compare the assumptions, which we supply for convenience in the Appendix. Firstly, Keribin (2000) shares A1–A3, and makes the additional condition (Id) requiring identifiability regarding mappings between the functional spaces \(\mathcal{M}_{k}^{\phi }\) and the parameter spaces \(\mathbb {S}_{k}\), which is unnecessary in our approach. Condition (P1-a) then requires that there exists a \(G\in \mathcal{L}_{1}\left( P\right) \) such that \(\left| \log f\right| \le G\), for every \(f\in \mathcal{M}_{\bar{k}}^{\phi }\), which is implied by A4. Condition (P1-b) then requires that the component density \(\phi \) possesses partial derivatives up to order 5, where all the partial derivatives of orders 2, 3, and 5, are each enveloped over \(\mathbb {T}\) by integrable functions in \(\mathcal{L}_{3}\left( \mathfrak {m}\right) \). This compares directly with A5, where we assume that \(\phi \) need not even have one derivative, only requiring that it is Lipschitz, with Lipschitz constant in \(\mathcal{L}_{1}\left( \mathfrak {m}\right) \) and that \(\phi \) itself is enveloped by an \(\mathcal{L}_{1}\left( \mathfrak {m}\right) \) function. This substantially simplifies the assumptions and makes our approach useful in many situations where the method of Keribin (2000) does not apply.

Conditions (P2) and (P3), which our theory does not require, then make linear independence assumptions between the first and second order partial derivatives of \(\phi \), and require that a class of linear combinations of zeroth- and first-order derivatives of \(\phi \) be a Donsker class of functionals with continuous sample paths, respectively. Finally, in the average-risk normalisation used here, Condition (C1) requires positive, increasing, vanishing penalties and imposes the true-order separation condition

$$ \lim _{n\rightarrow \infty }n\left\{ \textrm{pen}_{k,n}-\textrm{pen}_{k_{0},n}\right\} =\infty , $$

for each \(k_{0}<k\le \bar{k}\). Thus B2 strengthens the separation scale in Condition (C1) by a logarithmic factor. We believe this is a small price to pay for the reduction or elimination of the other conditions. In this setting we only require an infinitesimally larger penalty than that of the BIC, which is negligible in practice.

A complementary comparison can be made to the PanIC framework of Nguyen (2024), which treats model selection problems in a general empirical-risk minimisation form. For convenience, we reproduce in the Appendix the assumptions (N-A1)–(N-A2) and the penalty conditions (N-B1)–(N-B2) of Nguyen (2024). Specialising their setting to likelihood-based order selection by identifying \(\mathbb {T}_{k}\) with \(\mathbb {S}_{k}\) and taking the loss to be \(\ell _{k}(x;\psi _{k})=-\log f_{k}(x;\psi _{k})\), the hypotheses (N-A1)–(N-A2) amount to compactness of \(\mathbb {S}_{k}\) together with a global Lipschitz condition on the negative log-likelihood with an \(\mathcal{L}_{2}(P)\) envelope, which are slightly simpler to verify than A2–A5.

Under (N-A1)–(N-A2), Nguyen (2024, Thm. 1) establishes consistency for any penalised estimator whose penalties satisfy (N-B1)–(N-B2), where (N-B2) requires penalty size condition \(\sqrt{n}\{\textrm{pen}_{l,n}-\textrm{pen}_{k,n}\}\rightarrow \infty \) for \(k<l\), implying that \(\textrm{pen}_{l,n}-\textrm{pen}_{k,n}\) must dominate \(O\left( n^{-1/2}\right) \). By contrast, in the finite mixture likelihood setting, Lemma 4 shows that the relevant overfitting fluctuations are of order \(O\left( \log n/n\right) \). Consequently, the separation condition B2 in Theorem 1 only requires \(\textrm{pen}_{l,n}-\textrm{pen}_{k,n}\) to dominate \(O\left( \log n/n\right) \), permitting penalties that are asymptotically much smaller than those required by (N-B2) while delivering the same consistency guarantee that \(\textrm{P}(\hat{k}_{n}=k_{0})\rightarrow 1\) under A1–A5.

Remark 3

As discussed in Sect. 1, the penalty induced by the \(\epsilon \)-BIC and \(\nu \)-BIC can be made indistinguishable from that of the BIC when \(\epsilon \) is small or when \(\nu \) is large. In Table 1, we provide a schedule for the ratio of the rates of the \(\epsilon /\nu \)-BIC to that of the BIC, computed as \(\textrm{pen}_{k,n}^{\epsilon }/\textrm{pen}_{k,n}^{\textrm{BIC}}\) or \(\textrm{pen}_{k,n}^{\nu }/\textrm{pen}_{k,n}^{\textrm{BIC}}\), for the same choice of \(\alpha \), where \(\textrm{pen}_{k,n}^{\textrm{BIC}}=\alpha \left( k\right) n^{-1}\log n\). We observe that for sufficiently small choices of \(\epsilon \) or large choices of \(\nu \), the penalty ratios can be made to stay close to one for a larger range of sample sizes n.

Table 1 Schedule of penalty ratios \(\textrm{pen}_{k,n}^{\epsilon }/\textrm{pen}_{k,n}^{\textrm{BIC}}\) or \(\textrm{pen}_{k,n}^{\nu }/\textrm{pen}_{k,n}^{\textrm{BIC}}\) for the same choice of \(\alpha \)

Full size table

Remark 4

It would appear that to improve upon the fluctuation bound \(O\left( \log n/n\right) \) in Lemma 4 (equivalently, the normalising rate \(n/\log n\), and subsequently B2), sharper bounds on the entropy \(H_{\left[ \right] }\left( u,\bar{\mathcal{F}}^{1/2}\left( \delta \right) ,\mathcal{L}_{2}\left( \mathfrak {m}\right) \right) \) are required that make use of the localising condition \(\mathfrak {h}\left( f,f_{0}\right) \le \delta \) of \(\bar{\mathcal{F}}^{1/2}\left( \delta \right) \). It is in the effort of obtaining such sharp bounds that necessitates the additional assumptions regarding the partial derivatives of \(\phi \) employed by Keribin (2000) and Gassiat and Van Handel (2012). Indeed, van de Geer (2000, Sec. 7.5) notes that in the finite dimensional parametric setting, one cannot obtain \(O\left( n\right) \) rates without exploiting the localising condition. On the other hand, in Theorem 1, the penalty need only satisfy B1 and B2, and thus need only go to zero and scale at a rate larger than \(O\left( n^{-1}\log n\right) \). In fact the two choices that define the \(\nu \)-BIC and \(\epsilon \)-BIC are arbitrary and were chosen because they can be made arbitrarily close to the BIC in practice. However, one may choose any sufficiently large penalty. For example, in Nguyen (2024) the choice of \(\textrm{pen}_{k,n}=O\left( n^{-1/2}\sqrt{\textrm{Ln}^{\circ \nu }\left( n\right) }\right) \) is proposed, for \(\nu \ge 1\), in conforming with the suggestion in the treatment of Sin and White (1996), and applied successfully to mixture model settings in Nguyen (2024) and Westerhout et al. (2024), within the PanIC framework.

Remark 5

In Corollary 1, the choice of \(\alpha \) need only be a strictly increasing function in k and as such any such choice is valid for the conclusion to hold. For example, one may ignore the multiplier due to m and simply take \(\alpha \left( k\right) =k\), or one may take \(\alpha \left( k\right) \) to be proportional to the number of free parameters in \(\mathcal{M}_{k}^{\phi }\), as in the usual BIC. Concretely, since \(\psi _{k}\) contains \((k-1)\) free mixing weights and km component parameters, a natural choice is

$$ \alpha \left( k\right) =\frac{(m+1)k-1}{2}, $$

or the simplified version \(\alpha \left( k\right) =(m+1)k/2\) which we propose throughout, which ignores the simplex constraint. This indifference to the count of the number of so-called free parameters shows that one need not be so judicious regarding what defines a parameter, as in Fraley and Raftery (2002) and Baudry et al. (2010). Other admissible choices include \(\alpha \left( k\right) =k\) (counting mixture weights only; as considered in Keribin (2000)), \(\alpha \left( k\right) =km/2\) (counting component parameters only), or more aggressively \(\alpha \left( k\right) =k\log k\) or \(\alpha \left( k\right) =k^{\gamma }\) for \(\gamma >1\). Such freedom mirrors that in other consistent IC frameworks; for instance, the deterministic SWIC construction of Nguyen (2024, Rem. 3) permits any strictly increasing positive sequence to weight the common rate. More generally, the PanIC definition allows random penalty arrays, provided the corresponding stochastic penalty conditions are satisfied.

2.5 Misspecification

The results of Sect. 2.4 depend crucially on Assumption A1, which requires that \(f_{0}\in \mathcal{M}_{k_{0}}^{\phi }\) for some \(k_{0}\in \left[ \bar{k}\right] \). In practice, however, it is not always known that \(f_{0}\) belongs to the family of mixtures under study (and in particular, a “true order” \(k_{0}\) need not exist). In such a case, the conclusions of Theorem 1 and Corollary 1 are no longer meaningful as statements about recovering \(k_{0}\), but the information criteria considered above still admit a sensible population interpretation. We therefore replace A1 by the following assumption.

A1*:

\(X\in \mathbb {X}\) has PDF \(f_{0}\) with respect to \(\mathfrak {m}\), and \(\textbf{X}_{n}=\left( X_{1},\dots ,X_{n}\right) \) is a sample of IID replicates of X.

Recall the minimized risks \(L_{k}\) and \(\hat{L}_{k,n}\) from Sect. 2.1. Define the set of population risk minimising orders by

$$ \mathcal{K}^{*}=\arg \min _{k\in \left[ \bar{k}\right] }L_{k}. $$

Moreover, for any \(\psi _{k}\in \mathbb {S}_{k}\), whenever \(\textrm{E}\left\{ \left| \log f_{0}\left( X\right) \right| \right\} <\infty \), we have

$$ \ell _{k}\left( \psi _{k}\right) =-\textrm{E}\left\{ \log f_{0}\left( X\right) \right\} +\mathfrak {K}\!\left( f_{0}\Vert f_{k}\left( \cdot ;\psi _{k}\right) \right) , $$

and therefore minimising \(L_{k}\) over k is equivalent to minimising the Kullback–Leibler divergence \(\min _{\psi _{k}\in \mathbb {S}_{k}}\mathfrak {K}\!\left( f_{0}\Vert f_{k}\left( \cdot ;\psi _{k}\right) \right) \) over k. Equivalently,

$$ \mathcal{K}^{*}=\arg \min _{k\in \left[ \bar{k}\right] }\left\{ \min _{\psi _{k}\in \mathbb {S}_{k}}\mathfrak {K}\!\left( f_{0}\Vert f_{k}\left( \cdot ;\psi _{k}\right) \right) \right\} , $$

the set of KL minimising models among \(\left\{ \mathcal{M}_{k}^{\phi }\right\} _{k\in \left[ \bar{k}\right] }\). Without any further assumptions, we have the following result.

Proposition 2

Under A1*, A2–A4, and B1, \(\lim _{n\rightarrow \infty }\textrm{P}\!\left( \hat{k}_{n}\in \mathcal{K}^{*}\right) =1\).

Proposition 2 implies that under misspecification, any IC that satisfies B1 (and thus also any IC that satisfies B1 and B2, including the \(\epsilon \)-BIC and \(\nu \)-BIC) will eventually select an order whose best-fitting mixture model is KL-optimal among the candidate orders in \([\bar{k}]\). Besides the ICs studied in Sect. 2.4, many common ICs satisfy B1, including the AIC and the BIC. Consequently, if A1* and A2–A4 are satisfied, then all of these ICs admit the conclusion of Proposition 2.

The difference between Proposition 2 and Theorem 1 is that Proposition 2 only concludes that \(\hat{k}_{n}\) lies in the set \(\mathcal{K}^{*}\) with high probability, rather than that it converges to the smallest KL-minimising order (i.e., it is not a guarantee of parsimonious order selection). If one wishes to recover parsimonious order selection in the misspecified case, such a guarantee can be obtained via the PanIC framework of Nguyen (2024) by imposing the stronger penalty separation assumption (N-B2) from the Appendix; for instance, one may use a penalty of the form

$$ \textrm{pen}_{k,n}=\alpha \left( k\right) n^{-1/2}\textrm{Ln}^{\circ \nu }\left( n\right) , $$

with increasing function \(\alpha \).

We are not aware of general results proving parsimonious selection under misspecification for BIC-scale penalties, including the classical BIC, nor for our larger penalties satisfying B2, in finite mixture settings. In particular, Keribin (2000); Gassiat and Van Handel (2012), and Gassiat (2018, Sec. 4.3) all assume correct specification in their derivations. In Patilea (2001), the MLE theory of van de Geer (2000, Ch. 7) is developed for misspecified models. These results suggest a possible route to extending parsimonious order selection to misspecified settings, but carrying out such a program in the present finite-mixture framework would be nontrivial and goes beyond the scope of this work.

3 Example applications

In this section, we provide example applications of Theorem 1 and Corollary 1 via illustrations of how the sufficient conditions are verified.

3.1 Gaussian mixture models

We first make our theory concrete by considering the ubiquitous Gaussian mixture model, defined by taking

$$ \phi \left( x;\theta \right) =\textrm{det}\left( 2\pi \varSigma \right) ^{-1/2}\exp \left\{ -\frac{1}{2}\left( x-\mu \right) ^{\top }\varSigma ^{-1}\left( x-\mu \right) \right\} , $$

with \(x\in \mathbb {X}=\mathbb {R}^{p}\) and \(\theta =\left( \mu ,\varSigma \right) \in \mathbb {T}\), where \(\mathbb {T}\) is a compact subset of \(\mathbb {R}^{p}\times \mathbb {S}_{p}^{++}\). Here, \(\mathbb {S}_{p}^{++}\) denotes the set of positive definite matrices in \(\mathbb {R}^{p\times p}\) and \(\left( \cdot \right) ^{\top }\) denotes matrix transposition. Such a compact \(\mathbb {T}\) can be constructed, as per Ritter (2014, Sec. B.23), by taking

$$ \mathbb {T}=\left\{ \left( \mu ,\varSigma \right) \in \mathbb {R}^{p}\times \mathbb {S}_{p}^{++}:\left\| \mu \right\| \le b,c^{-1}\le \lambda _{1}\left( \varSigma \right) ,\lambda _{p}\left( \varSigma \right) \le c\right\} , $$

where \(\lambda _{1}\left( \varSigma \right) \) and \(\lambda _{p}\left( \varSigma \right) \) are the smallest and largest eigenvalues of \(\varSigma \), respectively, and \(b\ge 0\) and \(c\ge 1\) are fixed constants. By this construction, A2 and A3 are naturally verified.

Next, we can write

$$ \log \phi \left( x;\theta \right) =-\frac{1}{2}\log \textrm{det}\left( 2\pi \varSigma \right) -\frac{1}{2}\textrm{tr}\left\{ \left( x-\mu \right) \left( x-\mu \right) ^{\top }\varSigma ^{-1}\right\} , $$

and thus, with \(\textrm{E}\left\| X\right\| ^{2}<\infty \), we can verify A4 by taking

$$ G_{1}\left( x\right) =\frac{1}{2}\max _{\theta \in \mathbb {T}}\left| \log \textrm{det}\left( 2\pi \varSigma \right) \right| +\frac{1}{2}\max _{\theta \in \mathbb {T}}\left| \textrm{tr}\left\{ \left( x-\mu \right) \left( x-\mu \right) ^{\top }\varSigma ^{-1}\right\} \right| . $$

To verify A5*, we use standard expressions for the gradient of the normal density with respect to \(\theta \) (see, e.g., Loos 2016, Sec. 3.4.2). Write \(\theta =(\mu ,\varSigma )\), and let \(b\ge 0\) and \(c\ge 1\) be as in the definition of \(\mathbb {T}\). Define \(r(x)=\left[ \left\| x\right\| -b\right] _{+}\). Since \(\lambda _{1}(\varSigma )\ge c^{-1}\) and \(\lambda _{p}(\varSigma )\le c\), we have \(\textrm{det}(\varSigma )\ge c^{-p}\) and \(\varSigma ^{-1}\succeq c^{-1}I\), and therefore, for every \(x\in \mathbb {R}^{p}\) and \(\theta \in \mathbb {T}\),

$$ \begin{aligned} \phi (x;\theta )&\le (2\pi )^{-p/2}c^{p/2}\exp \left\{ -\frac{1}{2c}\Vert x-\mu \Vert ^{2}\right\} \\&\le (2\pi )^{-p/2}c^{p/2}\exp \left\{ -\frac{1}{2c}r(x)^{2}\right\} . \end{aligned} $$

Moreover, for the derivatives with respect to \(\mu \) and \(\varSigma \), we have

$$ \frac{\partial \phi \left( x;\theta \right) }{\partial \mu }=\varSigma ^{-1}\left( x-\mu \right) \phi \left( x;\theta \right) , $$

so that \(\left\| \partial \phi (x;\theta )/\partial \mu \right\| \le c\left\| x-\mu \right\| \phi (x;\theta )\). Similarly, the partial derivatives with respect to the coordinates of \(\varSigma \) can be written as \(\phi (x;\theta )\) multiplied by linear combinations of entries of \(\varSigma ^{-1}\) and of \(\varSigma ^{-1}(x-\mu )(x-\mu )^{\top }\varSigma ^{-1}\), and hence there exists a constant \(C_{0}>0\) depending only on b, c, and p, such that, for all x and \(\theta \in \mathbb {T}\),

$$ \max _{j\in \left[ m\right] }\left| \frac{\partial \phi \left( x;\theta \right) }{\partial \vartheta _{j}}\right| \le C_{0}\left\{ 1+\left\| x-\mu \right\| ^{2}\right\} \phi \left( x;\theta \right) . $$

Combining the preceding results and using \(\left\| x-\mu \right\| \le \left\| x\right\| +b\), we obtain

$$ \max _{\theta \in \mathbb {T}}\phi \left( x;\theta \right) +\max _{j\in \left[ m\right] }\max _{\theta \in \mathbb {T}}\left| \frac{\partial \phi \left( x;\theta \right) }{\partial \vartheta _{j}}\right| \le C_{1}\left\{ 1+\left\| x\right\| ^{2}\right\} \exp \left\{ -\frac{1}{2c}r(x)^{2}\right\} , $$

for a constant \(C_{1}>0\) depending only on b, c, and p. Therefore, A5* is satisfied by taking

$$ G_{2}\left( x\right) =C_{1}\left\{ 1+\left\| x\right\| ^{2}\right\} \exp \left\{ -\frac{1}{2c}r(x)^{2}\right\} , $$

noting that \(G_{2}\in \mathcal{L}_{1}\left( \mathfrak {m}\right) \) since it has a Gaussian tail. Since \(\textrm{E}\left\| X\right\| ^{2}<\infty \) holds under A1 for Gaussian mixture models, the conclusion of Theorem 1 holds for Gaussian mixture models whenever A1 holds. This is the archetypal example of Keribin (2000) and verifies the assumptions in the text.

3.2 Laplace mixture models

The Laplace mixture model is defined by taking the density of the Laplace distribution:

$$\begin{aligned} \phi \left( x;\theta \right) =\frac{\gamma }{2}\exp \left\{ -\gamma \left| x-\mu \right| \right\} , \end{aligned}$$

(5)

with \(x\in \mathbb {R}\) and \(\theta =\left( \mu ,\gamma \right) \in \mathbb {T}\), where we can choose

$$ \mathbb {T}=\left\{ \left( \mu ,\gamma \right) \in \mathbb {R}\times \mathbb {R}_{>0}:\left| \mu \right| \le b,c^{-1}\le \gamma \le c\right\} , $$

which is convex and compact for each \(b>0\) and \(c>1\). Such models have been considered, for example, by Cord et al. (2006); Mitianoudis and Stathaki (2007), and Rabbani and Vafadust (2008). Here, A2 and A3 hold by construction, although it is noteworthy that (5) is a nondifferentiable function of \(\mu \).

Similarly to Sect. 3.1, we can take

$$ G_{1}\left( x\right) =\max _{\theta \in \mathbb {T}}\left| \log \left( \frac{\gamma }{2}\right) \right| +\max _{\theta \in \mathbb {T}}\left\{ \gamma \left| x-\mu \right| \right\} $$

to verify A4, under the condition that \(\textrm{E}\left| X\right| <\infty \). Since \(\phi \left( x;\theta \right) \) is not differentiable in \(\mu \), we seek to verify A5 instead of A5*. Write \(\theta =(\mu ,\gamma )\) and let \(b>0\) and \(c>1\) be as in the definition of \(\mathbb {T}\). Define \(r(x)=\left[ |x|-b\right] _{+}\). Then, for every \(x\in \mathbb {R}\) and \(\theta \in \mathbb {T}\), we have \(|x-\mu |\ge r(x)\) and \(\gamma \ge c^{-1}\), so that

$$ \phi \left( x;\theta \right) =\frac{\gamma }{2}\exp \left\{ -\gamma |x-\mu |\right\} \le \frac{c}{2}\exp \left\{ -\frac{1}{c}r(x)\right\} . $$

To bound differences in \(\mu \), fix \(\gamma \in [c^{-1},c]\) and note that \(\bigl ||x-\mu _{1}|-|x-\mu _{2}|\bigr |\le |\mu _{1}-\mu _{2}|\). Using the mean value theorem for \(t\mapsto \exp \{-\gamma t\}\), we obtain

$$ \begin{aligned}&\left| \exp \{-\gamma |x-\mu _{1}|\}-\exp \{-\gamma |x-\mu _{2}|\}\right| \\&\qquad \le \gamma |\mu _{1}-\mu _{2}|\exp \{-\gamma r(x)\} \le c|\mu _{1}-\mu _{2}|\exp \left\{ -\frac{1}{c}r(x)\right\} , \end{aligned} $$

and therefore

$$ \left| \phi \left( x;(\mu _{1},\gamma )\right) -\phi \left( x;(\mu _{2},\gamma )\right) \right| \le \frac{c^{2}}{2}|\mu _{1}-\mu _{2}|\exp \left\{ -\frac{1}{c}r(x)\right\} . $$

To bound differences in \(\gamma \), fix \(\mu \in [-b,b]\) and use the mean value theorem in \(\gamma \). Since

$$ \frac{\partial }{\partial \gamma }\phi \left( x;(\mu ,\gamma )\right) =\frac{1}{2}\left\{ 1-\gamma |x-\mu |\right\} \exp \left\{ -\gamma |x-\mu |\right\} , $$

we have

$$ \begin{aligned} \left| \frac{\partial }{\partial \gamma }\phi (x;(\mu ,\gamma ))\right|&\le \frac{1}{2}\{1+\gamma |x-\mu |\}\exp \{-\gamma |x-\mu |\}\\&\le \frac{1}{2}\{1+c(|x|+b)\}\exp \left\{ -\frac{1}{c}r(x)\right\} . \end{aligned} $$

Combining the preceding bounds yields, for all \(\theta _{1},\theta _{2}\in \mathbb {T}\),

$$ \left| \phi \left( x;\theta _{1}\right) -\phi \left( x;\theta _{2}\right) \right| \le L_{\phi }\left( x\right) \left\| \theta _{1}-\theta _{2}\right\| _{1}, $$

where we can take

$$ L_{\phi }\left( x\right) =C_{0}\left\{ 1+|x|\right\} \exp \left\{ -\frac{1}{c}r(x)\right\} $$

for a suitable constant \(C_{0}>0\) depending only on b and c. Moreover, the bound on \(\phi (x;\theta )\) above shows that \(\max _{\theta \in \mathbb {T}}\phi (x;\theta )\le L_{\phi }(x)\) for \(C_{0}\) sufficiently large, and thus A5 is satisfied by taking \(G_{2}(x)=2L_{\phi }(x)\). Since \(G_{2}\in \mathcal{L}_{1}\left( \mathfrak {m}\right) \) by the exponential tails, we have the conclusion of Theorem 1 for Laplace mixture models whenever A1 is satisfied.

In this example, we observe that since \(\phi \) is not differentiable for all \(\theta \in \mathbb {T}\), the conditions (P1-b), (P2), and (P3) of Keribin (2000), defined via the partial derivatives of \(\phi \), cannot be verified.

3.3 Student t mixture models

The Student t mixture model is defined by taking the density of the Student t distribution:

$$\begin{aligned} \phi (x;\theta )=c_{v}\gamma \left( 1+\frac{\gamma ^{2}(x-\mu )^{2}}{v}\right) ^{-\frac{v+1}{2}}, \end{aligned}$$

(6)

with \(x\in \mathbb {R}\) and \(\theta =(\mu ,\gamma ,v)\in \mathbb {T}\), where

$$ c_{v}=\frac{\varGamma \left( \frac{v+1}{2}\right) }{\varGamma \left( \frac{v}{2}\right) \sqrt{v\pi }}. $$

We can choose

$$ \mathbb {T}=\left\{ (\mu ,\gamma ,v)\in \mathbb {R}\times \mathbb {R}_{>0}\times \mathbb {R}_{>0}:|\mu |\le b,\ c^{-1}\le \gamma \le c,\ \underline{v}\le v\le \overline{v}\right\} , $$

which is compact for each \(b>0\), \(c>1\), and \(0<\underline{v}<\overline{v}<\infty \). Mixtures of such distributions are considered in the works of Peel and McLachlan (2000b) and Forbes and Wraith (2014), among others.

Here, A2 and A3 hold by construction, and (6) is continuously differentiable in \(\theta \) on \(\mathbb {T}\). Similarly to Sect. 3.1, we can write

$$ \log \phi (x;\theta )=\log c_{v}+\log \gamma -\frac{v+1}{2}\log \!\left( 1+\frac{\gamma ^{2}(x-\mu )^{2}}{v}\right) , $$

and thus we can take

$$ G_{1}(x)=\max _{\theta \in \mathbb {T}}\left| \log (c_{v}\gamma )\right| +\frac{\overline{v}+1}{2}\log \!\left( 1+\frac{\bar{c}^{2}(|x|+b)^{2}}{\underline{v}}\right) ,\qquad \bar{c}=\max \{1/c,c\}, $$

to verify A4, under the condition that \(\textrm{E}X^{2}<\infty \) (using \(\log (1+u)\le u\) and \(|x-\mu |\le |x|+b\)).

Since \(\phi (x;\theta )\) is differentiable in \(\theta \), we seek to verify A5\(^{*}\). Define \(r(x)=\left[ |x|-b\right] _{+}\), so that \(|x-\mu |\ge r(x)\) for all \(|\mu |\le b\). First, since \(\gamma \le c\) and \(v\in [\underline{v},\overline{v}]\), we have \(\max _{\theta \in \mathbb {T}}(c_{v}\gamma )<\infty \). Moreover, since \(\gamma \ge c^{-1}\) and \(v\le \overline{v}\), we have

$$ 1+\frac{\gamma ^{2}(x-\mu )^{2}}{v}\ge 1+\frac{r(x)^{2}}{c^{2}\overline{v}}. $$

Since \(v\ge \underline{v}\), the map \(t\mapsto t^{-(v+1)/2}\) is bounded above by \(t^{-(\underline{v}+1)/2}\) for \(t\ge 1\), and hence there exists a constant \(C_{0}>0\) depending only on b, c, \(\underline{v}\), and \(\overline{v}\), such that

$$ \max _{\theta \in \mathbb {T}}\phi (x;\theta )\le C_{0}\left( 1+\frac{r(x)^{2}}{c^{2}\overline{v}}\right) ^{-\frac{\underline{v}+1}{2}}. $$

Next, for the derivatives, we recall that for \(\theta =(\mu ,\gamma ,v)\),

$$ \begin{aligned} \frac{\partial \phi (x;\theta )}{\partial \mu }&=(v+1)\frac{\gamma ^{2}(x-\mu )}{v+\gamma ^{2}(x-\mu )^{2}}\phi (x;\theta ),\\ \frac{\partial \phi (x;\theta )}{\partial \gamma }&=\left\{ \frac{1}{\gamma }-(v+1)\frac{\gamma (x-\mu )^{2}}{v+\gamma ^{2}(x-\mu )^{2}}\right\} \phi (x;\theta ), \end{aligned} $$

and

$$ \begin{aligned} \frac{\partial \phi (x;\theta )}{\partial v}=\phi (x;\theta )\Biggl \{&\frac{\partial }{\partial v}\log c_{v}-\frac{1}{2}\log \left( 1+\frac{\gamma ^{2}(x-\mu )^{2}}{v}\right) \\&+\frac{v+1}{2}\frac{\gamma ^{2}(x-\mu )^{2}}{v\{v+\gamma ^{2}(x-\mu )^{2}\}}\Biggr \}. \end{aligned} $$

Since \(t\mapsto t/(v+t^{2})\) is maximised at \(t=\sqrt{v}\) and \(\gamma \in [c^{-1},c]\), we have \(\gamma ^{2}|x-\mu |/\{v+\gamma ^{2}(x-\mu )^{2}\}\le c/(2\sqrt{\underline{v}})\), and thus

$$ \left| \frac{\partial \phi (x;\theta )}{\partial \mu }\right| \le C_{1}\phi (x;\theta ),\qquad \left| \frac{\partial \phi (x;\theta )}{\partial \gamma }\right| \le C_{2}\phi (x;\theta ), $$

for constants \(C_{1},C_{2}>0\) depending only on c, \(\underline{v}\), and \(\overline{v}\). Furthermore, \(\max _{v\in [\underline{v},\overline{v}]} |\partial _{v}\log c_{v}|<\infty \) and \(\max _{v\in [\underline{v},\overline{v}]}(v+1)/(2v)<\infty \), and therefore there exists a constant \(C_{3}>0\) such that

$$ \left| \frac{\partial \phi (x;\theta )}{\partial v}\right| \le \left\{ C_{3}+\frac{1}{2}\log \left( 1+\frac{\bar{c}^{2}(|x|+b)^{2}}{\underline{v}}\right) \right\} \phi (x;\theta ), \quad \bar{c}=\max \{1/c,c\}. $$

Combining these bounds with the envelope for \(\max _{\theta }\phi (x;\theta )\) above, we obtain that

$$ \max _{\theta \in \mathbb {T}}\phi (x;\theta )+\max _{j\in \left[ m\right] }\max _{\theta \in \mathbb {T}}\left| \frac{\partial \phi (x;\theta )}{\partial \vartheta _{j}}\right| \le G_{2}(x), $$

where we can take

$$ G_{2}(x)=C_{4}\left\{ 1+\log \left( 1+\frac{\bar{c}^{2}(|x|+b)^{2}}{\underline{v}}\right) \right\} \left( 1+\frac{r(x)^{2}}{c^{2}\overline{v}}\right) ^{-\frac{\underline{v}+1}{2}} $$

for a suitable constant \(C_{4}>0\). Since \(\underline{v}>0\), the right-hand side has a polynomial tail of order \(|x|^{-(\underline{v}+1)}\) up to a logarithmic factor, and hence \(G_{2}\in \mathcal{L}_{1}\left( \mathfrak {m}\right) \). Therefore, A5\(^{*}\) is satisfied.

Since \(\textrm{E}X^{2}<\infty \) holds under A1 whenever the smallest degrees of freedom among the true components satisfies \(v_{0}>2\), we have the conclusion of Theorem 1 for Student t mixture models whenever A1 is satisfied with \(v_{0}>2\).

In this example, we observe that although \(\phi \) is smooth, allowing v to vary can prevent verification of the integrability requirement in (P1-b) of Keribin (2000) (see the Appendix). Let \(v_{0}\) denote the smallest degrees of freedom among the true components, so that \(f_{0}(x)\asymp |x|^{-(v_{0}+1)}\) as \(|x|\rightarrow \infty \), and fix any candidate component with degrees of freedom \(v\in [\underline{v},\overline{v}]\), for which \(\phi (x;\theta )\asymp |x|^{-(v+1)}\). Then

$$ \int _{-\infty }^{\infty }\left( \frac{\phi (x;\theta )}{f_{0}(x)}\right) ^{3}f_{0}(x)\,dx=\int _{-\infty }^{\infty }\frac{\phi (x;\theta )^{3}}{f_{0}(x)^{2}}\,dx\asymp \int _{-\infty }^{\infty }|x|^{-(3v-2v_{0}+1)}\,dx, $$

which diverges whenever \(v\le \frac{2}{3}v_{0}\). Moreover, for fixed v, the derivative of \(\phi (x;\theta )\) with respect to \(\gamma \) is asymptotically a nonzero constant multiple of \(\phi (x;\theta )\) as \(|x|\rightarrow \infty \), so the corresponding derivative ratio has the same integrability obstruction. Consequently, if the parameter set \(\mathbb {T}\) contains degrees of freedom \(v\le 2v_{0}/3\) (in particular, if \(\underline{v}\le 2v_{0}/3\)), then (P1-b) cannot be verified and the conditions of Keribin (2000) are not applicable in such t mixture settings. For instance, if \(v_{0}=6\) then any model class allowing \(v\le 4\) violates (P1-b).

3.4 Mixture of regression models

One can extend upon finite mixture models via the mixture of regression construction which was introduced in Quandt (1972) and subsequently studied, for example, in Jones and McLachlan (1992); Wedel and DeSarbo (1995); Naik et al. (2007); Hafidi and Mkhadri (2010); Song et al. (2014); Depraetere and Vandebroek (2014), and Hui et al. (2015), among many others. A good recent account of mixture regression models appears in Yao and Xiang (2024). Suppose that we can write \(X=\left( U,Y\right) \), with \(\left( U,Y\right) :\varOmega \rightarrow \mathbb {U}\times \mathbb {Y}=\mathbb {X}\) and that we can write the PDF of X in the form

$$\begin{aligned} \phi \left( x;\theta \right) =\rho \left( y|u;\theta \right) \varphi \left( u\right) , \end{aligned}$$

(7)

for each \(x=\left( u,y\right) \in \mathbb {X}\), where \(\rho :\mathbb {U}\times \mathbb {Y}\times \mathbb {T}\rightarrow \mathbb {R}_{\ge 0}\) characterises a conditional PDF in the sense that for every \(\left( u,\theta \right) \in \mathbb {U}\times \mathbb {T}\), \(\int _{\mathbb {Y}}\rho \left( y|u;\theta \right) \mathfrak {m}_{2}\left( \textrm{d}y\right) =1\), and \(\varphi \) is the generative PDF of U with respect to the dominating measure \(\mathfrak {m}_{1}\), where \(\mathfrak {m}=\mathfrak {m}_{1}\times \mathfrak {m}_{2}\). Note here that \(\varphi \) is not parameterised by \(\theta \), nor is it given any structure and is merely assumed to exist.

Naturally, when fitting mixture of regression models, one considers only the conditional part of (7). As such, with

$$ \begin{aligned} \mathcal{N}_{k}^{\phi }=\Bigg \{\,\mathbb {U}\times \mathbb {Y}\ni (u,y)\mapsto \mathfrak {f}_{k}(y|u;\psi _{k})&=\sum _{z=1}^{k}\pi _{z}\rho (y|u;\theta _{z}):\\&\theta _{z}\in \mathbb {T},\ \pi _{z}\in [0,1],\ \sum _{z=1}^{k}\pi _{z}=1,\ z\in [k] \,\Bigg \}, \end{aligned} $$

for each \(k\in \left[ \bar{k}\right] \), one seeks to identify the smallest \(k_{0}\in \left[ \bar{k}\right] \), for which \(f_{0}=\mathfrak {f}_{0}\varphi \), where \(\mathfrak {f}_{0}\in \mathcal{N}_{k_{0}}^{\phi }\), or more concisely and in conforming to Sect. 1, \(f_{0}\in \mathcal{M}_{k_{0}}^{\phi }\). To this end, the usual risk, the average negative log-densities \(\ell _{k}\left( \psi _{k}\right) =-\textrm{E}\left\{ \log f_{k}\left( X;\psi _{k}\right) \right\} \), for each \(k\in \left[ \bar{k}\right] \) with \(f_{k}\in \mathcal{M}_{k}^{\phi }\), is not estimated by the natural estimator \(\ell _{k,n}\left( \psi _{k}\right) \), but instead by the empirical average negative log-conditional likelihood

$$ \mathfrak {L}_{k,n}\left( \psi _{k}\right) =-\frac{1}{n}\sum _{i=1}^{n}\log \mathfrak {f}_{k}\left( Y_{i}|U_{i};\psi _{k}\right) , $$

where \(\mathfrak {f}_{k}\in \mathcal{N}_{k}^{\phi }\).

A useful fact is that, since \(f_{k}\left( u,y;\psi _{k}\right) =\mathfrak {f}_{k}\left( y|u;\psi _{k}\right) \varphi \left( u\right) \), we have, for each \(k\in \left[ \bar{k}\right] \) and \(\psi _{k}\in \mathbb {S}_{k}\),

$$ \begin{aligned} \ell _{k,n}(\psi _{k})&=-P_{n}\log f_{k}(\cdot ;\psi _{k})\\&=-P_{n}\log \mathfrak {f}_{k}(\cdot ;\psi _{k})-P_{n}\log \varphi \\&=\mathfrak {L}_{k,n}(\psi _{k})-\frac{1}{n}\sum _{i=1}^{n}\log \varphi (U_{i}). \end{aligned} $$

Taking expectations yields \(\textrm{E}\left\{ \mathfrak {L}_{k,n}\left( \psi _{k}\right) \right\} =\ell _{k}\left( \psi _{k}\right) +\textrm{E}\left\{ \log \varphi \left( U\right) \right\} \), where the \(\textrm{E}\left\{ \log \varphi \left( U\right) \right\} \) term is the same across all k and has no influence on comparisons between risks for different values of k or different models within \(\mathcal{M}_{k}^{\phi }\) or \(\mathcal{N}_{k}^{\phi }\), since \(\varphi \) has no parameters. Moreover, the sample-level identity above shows that \(\ell _{k,n}\left( \psi _{k}\right) \) and \(\mathfrak {L}_{k,n}\left( \psi _{k}\right) \) differ by an additive term that does not depend on k or \(\psi _{k}\). Consequently, Lemma 3 and Lemma 4 remain valid with \(\mathfrak {L}_{k,n}\left( \psi _{k}\right) \) in place of \(\ell _{k,n}\left( \psi _{k}\right) \).

Furthermore, since \(P_{n}\log f=P_{n}\log \mathfrak {f}+P_{n}\log \varphi \), maximisers of the joint likelihood over \(\mathcal{M}_{k}^{\phi }\) correspond to maximisers of the conditional likelihood over \(\mathcal{N}_{k}^{\phi }\). In particular, if

$$ \hat{f}_{k,n}\in \underset{f\in \mathcal{M}_{k}^{\phi }}{\arg \max }P_{n}\log f, $$

then there exists

$$ \hat{\mathfrak {f}}_{k,n}\in \underset{\mathfrak {f}\in \mathcal{N}_{k}^{\phi }}{\arg \max }P_{n}\log \mathfrak {f}, $$

such that \(\hat{f}_{k,n}=\hat{\mathfrak {f}}_{k,n}\varphi \). If, in addition, \(\varphi \left( u\right) >0\) for all \(u\in \mathbb {U}\), then \(\log (\hat{f}_{k,n}/f_{0})=\log (\hat{\mathfrak {f}}_{k,n}/\mathfrak {f}_{0})\), and Proposition 1 is directly applicable to the maximum conditional likelihood estimator. Together, these facts imply that we can consistently estimate \(k_{0}\) by the conditional version of \(\hat{k}_{n}\):

$$ \tilde{k}_{n}=\min \arg \min _{k\in [\bar{k}]}\left\{ \min _{\psi _{k}\in \mathbb {S}_{k}}\mathfrak {L}_{k,n}(\psi _{k})+\textrm{pen}_{k,n}\right\} . $$

To be precise, if \(\phi \) has form (7), and Assumptions A1–A5 hold for the corresponding joint mixture class \(\mathcal{M}_{k}^{\phi }\), then Theorem 1 holds with \(\hat{k}_{n}\) replaced by \(\tilde{k}_{n}\), with no additional assumptions beyond A1–A5. This example is general and includes many models considered in the literature, as noted at the start. In particular, under compactness restrictions on \(\mathbb {T}\) and mild integrability conditions with respect to \(\mathfrak {m}_{1}\), the verification of A4 and A5 (or A5\(^{*}\)) for regression kernels \(\rho \) proceeds exactly as in Sects. 3.2 and 3.3, with x replaced by y and with the regression location depending on u. It is possible to verify these assumptions in settings with well-behaved \(\phi \), using the results of Keribin (2000), such as in the case of normal mixture regression models (e.g., Quandt (1972) and Jones and McLachlan (1992)). However, like in the Laplace mixture example, the Laplace mixture of regression models, such as that of Song et al. (2014), have non-differentiable PDFs and thus violate the same conditions of Keribin (2000) as in the previous example. Mixture of Student t regression models such as those of Galimberti and Soffritti (2014) and Yao et al. (2014) will also incur the same problems as described in Sect. 3.3 if the degree of freedom parameters considered are too small.

Remark 6

Since \(\varphi \) is not parameterised by \(\theta \), it does not influence the maximisation of the likelihood over \(\psi _{k}\), and it cancels from likelihood ratios \(\log (\hat{f}_{k,n}/f_{0})\) whenever \(\varphi \left( u\right) >0\) for all \(u\in \mathbb {U}\). However, \(\varphi \) may enter the verification of A1–A5 through integrability requirements with respect to the dominating measure \(\mathfrak {m}=\mathfrak {m}_{1}\times \mathfrak {m}_{2}\). A convenient convention is to take \(\mathfrak {m}_{1}\) to be the marginal law of U, in which case \(\varphi =1\) and \(\phi \left( x;\theta \right) =\rho \left( y|u;\theta \right) \) is a density with respect to \(\mathfrak {m}=\mathfrak {m}_{1}\times \mathfrak {m}_{2}\). Under this convention, the verification of A1–A5 reduces to verifying them for \(\rho \).

3.5 Numerical simulations

Here, we provide some numerical evidence towards the comparative performance of the \(\epsilon \)-BIC and \(\nu \)-BIC versus alternatives such as the AIC, BIC, and PanIC. We consider four scenarios corresponding to our technical examples in Sects. 3.2 and 3.3, which present pathologies that do not verify the hypotheses of Keribin (2000). In each case, we simulate IID data \(\textbf{X}_{n}\) from the measure with density \(f_{0}\), a mixture with true number of components \(k_{0}\) of base density \(\phi \) with true parameter \(\psi _{k_{0}}^{*}\). Then, we estimate \(\hat{k}_{n}\) by considering mixtures of up to \(K>k_{0}\) components of the same form, using data of sizes \(n\in \left\{ 100,1000,10000\right\} \). The four scenarios considered are described in Table 2.

Table 2 Description of simulation scenarios

Full size table

We repeat each scenario 200 times, and evaluate the proportions of times \(\hat{k}_{n}=k\), for each \(k\in \left[ K\right] \) and for each IC. Each IC is considered to have penalties in the form \(\textrm{pen}_{k,n}=\alpha \left( k\right) r_{n}\), where \(r_{n}\) is the shape of the penalty. For the \(\epsilon \)-BIC, we consider \(\epsilon =0.1\) and \(\epsilon =0.02\). For the \(\nu \)-BIC, we consider \(\nu =1\) (also corresponding to the \(\epsilon \)-BIC with \(\epsilon =1\)), and \(\nu =3\). For PanIC, we use the Sin–White information criterion form considered in Nguyen (2024). The shapes for the considered penalties are thus: AIC: \(r_{n}=2n^{-1}\); BIC: \(r_{n}=n^{-1}\log n\); \(\epsilon \)-BIC: \(r_{n}=n^{-1}\log ^{1+\epsilon }n\); \(\nu \)-BIC: \(r_{n}=n^{-1}\textrm{Ln}^{\circ \nu }\left( n\right) \log n\); PanIC: \(r_{n}=n^{-1/2}\sqrt{\textrm{Ln}^{\circ 2}\left( n\right) }\). The constant multiples \(\alpha \left( k\right) \) that we consider are of forms (I) \(\textrm{dim}\left( \mathbb {S}_{k}\right) /2\) (the classic BIC penalty), (II) \(\left\{ \textrm{dim}\left( \mathbb {S}_{k}\right) \right\} ^{2}/2\), (III) k (only the number of mixture components), and (IV) \(k^{2}\). The results are recorded in Tables 3, 4, 5, 6.

Table 3 Results from 200 replications of scenario S-L1. The \(\hat{k}_{n}=\dots \) columns indicate the proportion of times \(\hat{k}_{n}=k\), where \(*\) marks the true number of components \(k_{0}\). Bold text indicates the highest proportion of correct identifications of \(k_{0}\)

Full size table

Table 4 Results from 200 replications of scenario S-L2. The \(\hat{k}_{n}=\dots \) columns indicate the proportion of times \(\hat{k}_{n}=k\), where \(*\) marks the true number of components \(k_{0}\). Bold text indicates the highest proportion of correct identifications of \(k_{0}\)

Full size table

Table 5 Results from 200 replications of scenario S-T1. The \(\hat{k}_{n}=\dots \) columns indicate the proportion of times \(\hat{k}_{n}=k\), where \(*\) marks the true number of components \(k_{0}\). Bold text indicates the highest proportion of correct identifications of \(k_{0}\)

Full size table

Table 6 Results from 200 replications of scenario S-T2. The \(\hat{k}_{n}=\dots \) columns indicate the proportion of times \(\hat{k}_{n}=k\), where \(*\) marks the true number of components \(k_{0}\). Bold text indicates the highest proportion of correct identifications of \(k_{0}\)

Full size table

From Tables 36, we observe, as expected, that as n increases, the proportions of times that the \(\epsilon \)-BIC, \(\nu \)-BIC, and PanIC estimate the correct number of components increase, and typically become perfectly accurate, except in cases where the penalty and shape combinations are too severe and thus consistency is not realised in practice even at \(n=10000\), such as for PanIC with Type II penalties or the \(\nu =1\) case of the \(\nu \)-BIC with Type II penalties. We observe in our experiments that the Type III penalties when applied with the \(\epsilon \)-BIC or \(\nu \)-BIC shapes tend to be too small to realise perfect accuracy at \(n=10000\), although it is clear that the proportion of times \(\hat{k}_{n}=k_{0}\) is increasing in each case. In terms of good choices for \(\alpha \), we observe that Type I and Type IV forms appear to work well with all penalty shapes, with Type II often being too large and Type III being too small.

Perhaps the most surprising outcome is that for some configurations, the AIC and the BIC can produce perfect or near perfect estimation of \(k_{0}\) for large n. Indeed, in the case of the AIC, as we shall show in Proposition 7 (albeit in a simplified setting that does not include the current experimental scenarios), it is impossible to guarantee that the proportion of accurate estimates of \(k_{0}\) goes to one. This points to either this being a small sample phenomenon and thus the true proportion of overestimating \(k_{0}\) is possibly small but non-zero, or that there exist settings where the AIC is consistent, although we believe that the former is more likely than the latter given our results in Sect. 4.2. In the case of the BIC, we argue that it is unknown whether the conditions of Keribin (2000) are necessary and sufficient, and that there is potentially a gap between our theory and the theory of Keribin (2000). Furthermore, if we observe the results of the \(\epsilon \)-BIC with \(\epsilon =0.02\) and the \(\nu \)-BIC with \(\nu =3\), we note that the outcomes are either exactly the same as or very close to those of the BIC for all n, due to the fact that, numerically, the criteria are only infinitesimally different for any practical size n. We, however, conclude that there is indeed evidence that the infinitesimal variations from the BIC, via the \(\epsilon \)-BIC and \(\nu \)-BIC constructions, are a harmless and practically costless way to guarantee consistency in settings where such a guarantee may not otherwise be afforded by the results of Keribin (2000). Overall, these simulations lead to our recommendation, at least in these simulation scenarios, that the \(\epsilon \)-BIC with \(\epsilon =0.02\) and the \(\nu \)-BIC with \(\nu =3\), together with \(\alpha \) of Type I or Type IV, provide good finite sample performance while enabling the guarantees of Corollary 1.

4 Limitations
4.1 Penalty flexibility

Theorem 1 provides a very general and flexible approach to order selection for mixture models, and as pointed out in Remarks 4 and 5, there are many possible choices for penalty calibration beyond our proposed \(\nu \)-BIC and \(\epsilon \)-BIC proposals. A natural question to ask is whether there are optimal choices for the penalty. In particular, from a minimax perspective (cf. Gassiat and Van Handel (2012) and Gassiat (2018), Ch. 4), we would like to determine if there is a smallest penalty that satisfies B1 and B2. Such a penalty would provide model selection consistency while minimally modifying the maximum likelihood as the axis for model choice.

Unfortunately, while such a minimal penalty is desirable, we can prove that it does not exist. In particular, for fixed \(\bar{k}\ge 2\), let \(\mathscr {P}\) denote the collection of penalty arrays \(\textrm{pen}=\left\{ \textrm{pen}_{k,n}:k\in \left[ \bar{k}\right] ,n\in \mathbb {N}\right\} \) that satisfy B1 and B2. We endow \(\mathscr {P}\) with the domination order: \(\textrm{pen}^{\left( 1\right) }\preceq \textrm{pen}^{\left( 2\right) }\) if and only if there exists an \(n_{0}\in \mathbb {N}\) such that for all \(n\ge n_{0}\) and for all \(k\in \left[ \bar{k}\right] \), \(\textrm{pen}_{k,n}^{\left( 1\right) }\le \textrm{pen}_{k,n}^{\left( 2\right) }\). Call \(\textrm{pen}^{*}\in \mathscr {P}\) universally optimal if it is a least element of \(\left( \mathscr {P},\preceq \right) \); i.e., \(\textrm{pen}^{*}\preceq \textrm{pen}\), for each \(\textrm{pen}\in \mathscr {P}\).

Proposition 3

The partial order \(\left( \mathscr {P},\preceq \right) \) has no minimal element.

More precisely, for every \(\textrm{pen}\in \mathscr {P}\), there exists a \(\widetilde{\textrm{pen}}\in \mathscr {P}\) such that \(\widetilde{\textrm{pen}}\preceq \textrm{pen}\) and such that the domination is strict for at least one order for all sufficiently large n; indeed, it is strict for every \(k\ge 2\) for all sufficiently large n. In particular, this shows that there is no optimal choice of penalty rate and \(\alpha \) function among the class of IC penalties characterised by B1 and B2 alone, and any penalty that satisfies B1 and B2 admits an asymptotically smaller penalty that also satisfies B1 and B2.

We may hope that if we have some fixed \(\alpha :\mathbb {N}\rightarrow \mathbb {R}_{>0}\) strictly increasing, and we define our penalty in the form \(\textrm{pen}_{k,n}^{\alpha }=\alpha \left( k\right) r_{n}\), where \(\left( r_{n}\right) _{n}\subset \mathbb {R}_{>0}\), then there may be some minimal choice of \(\left( r_{n}\right) _{n}\). Unfortunately, this also is not true as per the following result.

Proposition 4

For penalties of the form \(\textrm{pen}_{k,n}^{\alpha }=\alpha \left( k\right) r_{n}\), if B1 and B2 hold then necessarily \(r_{n}\rightarrow 0\) and \(\left( n/\log n\right) r_{n}\rightarrow \infty \), as \(n\rightarrow \infty \). Moreover, for each such \(\left( r_{n}\right) _{n}\), there exists another sequence \(\left( \tilde{r}_{n}\right) _{n}\) such that \(\tilde{r}_{n}=o\left( r_{n}\right) \), while still satisfying \(\tilde{r}_{n}\rightarrow 0\) and \(\left( n/\log n\right) \tilde{r}_{n}\rightarrow \infty \), as \(n\rightarrow \infty \).

Proof

Let \(a_{n}=\left( n/\log n\right) r_{n}\rightarrow \infty \) and set \(\tilde{r}_{n}=r_{n}/\sqrt{a_{n}}\). Then \(\tilde{r}_{n}/r_{n}=a_{n}^{-1/2}\rightarrow 0\), hence \(\tilde{r}_{n}=o\left( r_{n}\right) \) and \(\tilde{r}_{n}\rightarrow 0\). Lastly, \(\left( n/\log n\right) \tilde{r}_{n}=\sqrt{a_{n}}\rightarrow \infty \), satisfying B2, since \(\alpha \left( l\right) -\alpha \left( k\right) >0\) for each \(l>k\), as required. \(\square \)

Indeed, the same penalty flexibility issue is not only a feature of our consistency theorems but also of the PanIC approach of Nguyen (2024) and Westerhout et al. (2024), the preceding results of Sin and White (1996), and even the mixture model BIC consistency results of Keribin (2000); Gassiat and Van Handel (2012), and Gassiat (2018). This is also a problem in the finite sample model selection approaches of Birgé and Massart (2007) and Massart (2007) (implemented in the mixture model literature in the works of Maugis and Michel (2011); Maugis-Rabusseau and Michel (2013), and Nguyen et al. (2022)). In the finite sample setting, some methods are available for empirical calibration of penalties, such as via the so-called slope heuristic (cf. Baudry et al. (2012) and Arlot (2019)). Although there is some evidence that these methods work well in empirical studies, they provide no technical or concrete guarantees in the mixture model setting.

4.2 Inefficiency of consistent model selection

In Yang (2005), it was demonstrated that, in the regression context, there is a fundamental tension between the AIC and consistent estimators such as the BIC. In particular, it was shown that in the context of choosing between nested regression models, the chosen model selected via the AIC is minimax optimal in the mean squared error sense, while it is impossible for any fitted model chosen via a model consistent estimator to achieve the same minimax optimality. In this subsection, we will prove that the same phenomenon arises in the mixture model order selection setting.

To proceed, we will consider a simplified version of the example in Sect. 3.1. In particular, let

$$ x\mapsto \phi \left( x;\mu \right) =\frac{1}{\sqrt{2\pi }}\exp \left\{ -\frac{1}{2}\left( x-\mu \right) ^{2}\right\} $$

be the density of the univariate normal law with unit variance \(\textrm{N}\left( \mu ,1\right) \). For fixed \(b>0\) and \(\underline{\pi }\in \left( 0,1/2\right) \), let \(\mathcal{F}=\mathcal{M}_{1}\cup \mathcal{M}_{2}\), where \(\mathcal{M}_{1}=\left\{ \phi \left( \cdot ;\mu \right) :\mu \in \left[ -b,b\right] \right\} \), and

$$ \begin{aligned} \mathcal{M}_{2}=\{\,&\pi \phi (\cdot ;\mu _{1})+(1-\pi )\phi (\cdot ;\mu _{2}): \pi \in [\underline{\pi },1-\underline{\pi }],\\&\mu _{1},\mu _{2}\in [-b,b],\ \mu _{1}\le \mu _{2}\,\}. \end{aligned} $$

Let \(\textbf{X}_{n}\) consist of IID replicates from the law with density (with respect to the Lebesgue measure) \(f_{0}\in \mathcal{F}\). As in Sect. 3.4, we let \(\hat{f}_{k,n}\in \arg \max _{f\in \mathcal{M}_{k}}P_{n}\log f\) for \(k\in \left\{ 1,2\right\} \), and with \(f\mapsto \ell _{n}\left( f\right) =-n^{-1}\sum _{i=1}^{n}\log f\left( X_{i}\right) \), we write

$$ \begin{aligned} \hat{k}_{n}^{\textrm{AIC}}&=\min \underset{k\in \{1,2\}}{\arg \min }\left\{ \ell _{n}(\hat{f}_{k,n})+\textrm{pen}_{k,n}^{\textrm{AIC}}\right\} ,\\ \hat{f}_{n}^{\textrm{AIC}}&=\hat{f}_{\hat{k}_{n}^{\textrm{AIC}},n}. \end{aligned} $$

Here, we define AIC-like penalties via \(\textrm{pen}_{k,n}^{\textrm{AIC}}=a\textrm{dim}\left( \mathbb {S}_{k}\right) n^{-1}\), where \(a>0\) is sufficiently large, and \(\textrm{dim}\left( \mathbb {S}_{1}\right) =1\) and \(\textrm{dim}\left( \mathbb {S}_{2}\right) =3\), in this case. As per the usual treatment on the topic (see, e.g., Wainwright (2019), Ch. 15), we define the squared-Hellinger minimax risk as

$$ \mathscr {R}_{n}^{*}=\inf _{\tilde{f}_{n}}\sup _{f_{0}\in \mathcal{F}}\textrm{E}_{f_{0}}\left\{ \mathfrak {h}^{2}\left( \tilde{f}_{n},f_{0}\right) \right\} , $$

where \(\textrm{E}_{f_{0}}\) denotes the expectation over \(\textbf{X}_{n}\) with respect to the law with density \(f_{0}\). The infimum is taken over the class of all measurable estimators \(\tilde{f}_{n}\). The following result states that the AIC-like estimator \(\hat{f}_{n}^{\textrm{AIC}}\) achieves the minimax risk.

Proposition 5

For each sufficiently large \(a>0\), there exist constants \(0<c<C<\infty \) depending on a, b, and \(\underline{\pi }\) such that, for all sufficiently large n,

$$ \frac{c}{n}\le \mathscr {R}_{n}^{*}\le \sup _{f_{0}\in \mathcal{F}}\textrm{E}_{f_{0}}\left\{ \mathfrak {h}^{2}\left( \hat{f}_{n}^{\textrm{AIC}},f_{0}\right) \right\} \le \frac{C}{n}. $$

In contrast to the AIC-like penalties, we note that the BIC, and more generally any penalty satisfying B1 together with B2, verifies the condition

$$\begin{aligned} \varDelta _{n}\underset{n\rightarrow \infty }{\longrightarrow }0,\qquad n\varDelta _{n}\underset{n\rightarrow \infty }{\longrightarrow }\infty , \end{aligned}$$

(8)

where \(\varDelta _{n}=\textrm{pen}_{2,n}-\textrm{pen}_{1,n}\). The same implication holds for the average-risk form of Condition (C1) of Keribin (2000) and for penalties satisfying (N-B1)–(N-B2) of Nguyen (2024). With \(\tilde{f}_{n}=\hat{f}_{\hat{k}_{n},n}\), where \(\hat{k}_{n}=\arg \min _{k\in \left\{ 1,2\right\} }\left\{ \ell _{n}\left( \hat{f}_{k,n}\right) +\textrm{pen}_{k,n}\right\} \), we have the following result.

Proposition 6

If (8) holds then there exists a constant \(c>0\) depending only on b and \(\underline{\pi }\), such that for all sufficiently large n,

$$ \sup _{f_{0}\in \mathcal{F}}\textrm{E}_{f_{0}}\left\{ \mathfrak {h}^{2}\left( \tilde{f}_{n},f_{0}\right) \right\} \ge c\varDelta _{n}. $$

Proposition 6 implies that any selector \(\hat{k}_{n}\) that satisfies (8) cannot achieve the minimax Hellinger risk over \(\mathcal{F}\). Indeed, from Proposition 5, we have that \(\mathscr {R}_{n}^{*}\asymp n^{-1}\), which, combined with Proposition 6, yields

$$ \frac{\sup _{f_{0}\in \mathcal{F}}\textrm{E}_{f_{0}}\left\{ \mathfrak {h}^{2}\left( \tilde{f}_{n},f_{0}\right) \right\} }{\mathscr {R}_{n}^{*}}\gtrsim n\varDelta _{n}\rightarrow \infty . $$

Thus, the worst-case risk of \(\tilde{f}_{n}\) is asymptotically larger than the minimax risk by an unbounded factor. A natural follow-up question is whether it is possible to construct a consistent selector that does not verify condition (8). The following result demonstrates that this is impossible in the current setting.

Proposition 7

Suppose that \(\hat{k}_{n}\) is the penalised likelihood selector on \(\mathcal{F}=\mathcal{M}_{1}\cup \mathcal{M}_{2}\) with penalty gap \(\varDelta _{n}=\textrm{pen}_{2,n}-\textrm{pen}_{1,n}\). If \(\hat{k}_{n}\) is order-consistent in the sense that, for every \(f_{0}\in \mathcal{M}_{1}\), \(\textrm{P}_{f_{0}}(\hat{k}_{n}=2)\rightarrow 0\), and, for every \(f_{0}\in \mathcal{M}_{2}\setminus \mathcal{M}_{1}\), \(\textrm{P}_{f_{0}}(\hat{k}_{n}=1)\rightarrow 0\), then (8) necessarily holds.

5 Discussion

The IC approach to model selection and order selection is ubiquitous across many domains of statistical modeling, with particular popularity and applicability in the context of finite mixture models. Among the IC approaches, the BIC is often the most commonly used, and in the mixture model setting where the maximum model size \(\bar{k}\) is known, its consistency has been proved in Keribin (2000). In this work, we substantially relax the assumptions made by Keribin (2000) via minor modifications of the BIC, to produce the \(\nu \)-BIC and \(\epsilon \)-BIC. These ICs minimally modify the penalty term of the BIC but allow for much broader applicability, notably being applicable without any differentiability and requiring substantially weaker regularity than high-order differentiability and moment conditions. Via example applications, we show that these ICs are consistent estimators in the popular setting of Gaussian mixture models, and also in settings that fall outside the scope of Keribin (2000), including non-differentiable Laplace mixture models, heavy-tailed t-mixture models, and mixtures of regression models. In addition, we provide a complementary result under misspecification: when \(f_{0}\) need not belong to any \(\mathcal{M}_{k}^{\phi }\), any IC satisfying B1 will eventually select an order whose best-fitting mixture model is Kullback–Leibler optimal among the candidate orders in \([\bar{k}]\) (Proposition 2).

As we have discussed in Sect. 1, for suitable and reasonable choices of \(\nu \) and \(\epsilon \), the respective \(\nu \)-BIC and \(\epsilon \)-BIC are numerically indistinguishable from the BIC in practical settings. We accompany this observation with numerical simulations in Sect. 3, which compare the \(\epsilon \)-BIC and \(\nu \)-BIC against alternatives such as the AIC, BIC, and PanIC in scenarios that exhibit the pathologies highlighted in our technical examples. In these experiments, we observe that choices such as \(\epsilon =0.02\) and \(\nu =3\) are, for all practical values of n, effectively the same as the BIC numerically, while still inheriting the theoretical guarantees of Corollary 1 in settings where the hypotheses of Keribin (2000) are not available. More generally, the simulation study provides some empirical guidance on the practical calibration of \(\alpha \) and on the behaviour of competing criteria in difficult regimes.

Given the potentially small numerical difference between the \(\nu \)-BIC and \(\epsilon \)-BIC and the BIC, we do not seek to push the use of either criterion over the BIC, which many practitioners are already accustomed to and use regularly. However, we wish to position our results as technical tools for explaining why penalties that are numerically close to the BIC can remain valid even when the stringent assumptions of Keribin (2000) are not met. For example, one may use our theory to justify BIC-like modifications for mixtures with non-differentiable components \(\phi \), such as triangular mixtures (Karlis and Xekalaki, 2008; Nguyen and McLachlan, 2016) and Laplace-based mixture models (Franczak et al., 2013; Song et al., 2014; Azam and Bouguila, 2020), as well as mixtures with components \(\phi \) with limited regularity, such as t-distribution based models (Peel and McLachlan, 2000b; Lin et al., 2007; Lo and Gottardo, 2012; Forbes and Wraith, 2014; Yao et al., 2014). In the misspecified case, Proposition 2 further indicates that, under mild conditions, common ICs (including the AIC and the BIC) still target Kullback–Leibler optimal approximating models, albeit without a guarantee of parsimonious order selection.

Indeed, via Theorem 1, it is possible to construct multitudes of ICs that exhibit consistency in mixture model settings, but the finite sample performance of any such penalty cannot be deduced from our general consistency theory alone. In particular, as formalised in Sect. 4.2, there are two structural limitations that consistency results do not resolve. First, under the natural domination order, there is no universally minimal penalty array satisfying B1 and B2, so consistency does not single out an optimal BIC-scale penalty calibration. Second, we establish a tension between order consistency and minimax optimality in mixtures like that of Yang (2005): in a simple Gaussian mixture family, an AIC-like criterion achieves the parametric minimax Hellinger risk, whereas any order-consistent criterion must be minimax-inefficient by an unbounded factor, typically logarithmic for BIC-scale penalties. To the best of our knowledge, this AIC/BIC tension has not previously been stated explicitly in the finite-mixture order selection setting. In the finite sample setting, some methods are available for empirical calibration of penalties, such as via the so-called slope heuristic (cf. Baudry et al. (2012) and Arlot (2019)), although such methods provide no general guarantees in the mixture model setting.

Although the assumptions of Theorem 1 are weaker than those of Keribin (2000), it is still possible to deduce consistent model selection under further relaxations of conditions. For example, even without continuity of the component density \(\phi \) or independence of the data \(\textbf{X}_{n}\), or when the loss function is not necessarily the negative log-density, we can deduce consistency using the theory of Westerhout et al. (2024), at the expense of increasing the rate of the penalisation term to \(\tilde{O}\left( n^{-1/2}\right) \). An immediate direction of study is to understand whether further relaxations of the assumptions are possible without such a pronounced increase in the penalisation rate. Another natural direction, motivated by truth-changing and locally misspecified frameworks in modern model selection (for example the focused information criterion literature of Claeskens and Hjort (2003); Lohmeyer et al. (2019); Pandhare and Ramanathan (2020a, 2020b) and related high-dimensional settings such as Gao and Carroll (2017)), is to extend our results to triangular-array regimes where the effective data-generating mechanism, or the KL-optimal order, may vary with n. We expect such extensions to be possible, but they would require triangular-array analogues of the empirical process arguments underpinning our uniform convergence and overfitting rate results. In addition, under misspecification, Proposition 2 only ensures that \(\hat{k}_{n}\) eventually lies in the set of KL-minimising orders, and we are not aware of general results proving parsimonious order selection under misspecification for BIC-scale penalties. It would therefore be of interest to understand whether parsimonious selection is achievable under misspecification without moving to larger penalties, or whether such a limitation is inherent.

Another immediate extension to our current work is to obtain the corresponding regularity conditions for consistent order selection of mixture of experts models (cf. Jacobs et al. (1991); Yuksel et al. (2012), and Nguyen and Chamroukhi (2018)), which extend upon the finite mixture models by also parameterising the mixing coefficients \(\pi _{1},\dots ,\pi _{k}\). Such an extension would add to the growing body of theoretical results regarding model selection and model uncertainty in mixture of experts models, including the results of Khalili (2010); Nguyen et al. (2022, 2023c, 2024c, 2023, 2024a, 2024b); Thai et al. (2025), and Westerhout et al. (2024), for example. Lastly, a limitation of our current results is the requirement that some maximum order \(\bar{k}\) be specified. Indeed such an assumption can be omitted at the great expense of strong control of local bracketing entropy of normalisations of the mixture model classes, as in Gassiat and Van Handel (2012) and Gassiat (2018, Sec. 4.3). It remains open as to whether weaker assumptions are available, even for penalties with the larger rate \(\tilde{O}\left( n^{-1/2}\right) \).

References
  • Akaike, H. (1974). A new look at the statistical model identification. IEEE Transactions on Automatic Control, 19, 716–723.

    Article  MathSciNet  Google Scholar 

  • Anderson, D., Burnham, K. (2004). Model selection and multi-model inference. New York: Springer.

    Google Scholar 

  • Arlot, S. (2019). Minimal penalties and the slope heuristics: a survey. Journal de la Société Francaise de Statistique, 160, 1–106.

    MathSciNet  Google Scholar 

  • Azam, M., Bouguila, N. (2020). Multivariate bounded support laplace mixture model. Soft Computing, 24, 13239–13268.

    Article  Google Scholar 

  • Baudry, J.-P. (2015). Estimation and model selection for model-based clustering with the conditional classification likelihood. Electronic Journal of Statistics, 9(1), 1041–1077.

    Article  MathSciNet  Google Scholar 

  • Baudry, J.-P., Raftery, A. E., Celeux, G., Lo, K., Gottardo, R. (2010). Combining mixture components for clustering. Journal of Computational and Graphical Statistics, 19(2), 332–353.

    Article  MathSciNet  Google Scholar 

  • Baudry, J.-P., Maugis, C., Michel, B. (2012). Slope heuristics: overview and implementation. Statistics and Computing, 22, 455–470.

    Article  MathSciNet  Google Scholar 

  • Birgé, L., Massart, P. (2007). Minimal penalties for gaussian model selection. Probability theory and related fields, 138, 33–73.

    Article  MathSciNet  Google Scholar 

  • Budanova, S. (2025). Penalized estimation of finite mixture models. Journal of Econometrics, 249, 105958.

    Article  MathSciNet  Google Scholar 

  • Chen, J. (2023). Statistical inference under mixture models. Singapore: Springer.

    Book  Google Scholar 

  • Chen, J., Khalili, A. (2009). Order selection in finite mixture models with a nonsmooth penalty. Journal of the American Statistical Association, 104, 187–196.

    Article  MathSciNet  Google Scholar 

  • Chiu Chong, M., Nguyen, H. D., Nguyen, T. (2024). Risk bounds for mixture density estimation on compact domains via the h-lifted kullback–leibler divergence. Transactions on Machine Learning Research.

  • Claeskens, G., Hjort, N. L. (2003). The focused information criterion (With discussion). Journal of the American Statistical Association, 98(464), 900–916.

    Article  MathSciNet  Google Scholar 

  • Claeskens, G., Hjort, N. L. (2008). Model selection and model averaging. Cambridge: Cambridge University Press.

    Google Scholar 

  • Cord, A., Ambroise, C., Cocquerez, J.-P. (2006). Feature selection in robust clustering based on laplace mixture. Pattern Recognition Letters, 27, 627–635.

    Article  Google Scholar 

  • Cormen, T. H., Leiserson, C. E., Rivest, R. L., Stein, C. (2002). Introduction to algorithms. Cambridge, MA: MIT Press.

    Google Scholar 

  • Csiszár, I. (1995). Generalized projections for non-negative functions. Acta Mathematica Hungarica, 68, 161–186.

    Article  MathSciNet  Google Scholar 

  • Depraetere, N., Vandebroek, M. (2014). Order selection in finite mixtures of linear regressions: literature review and a simulation study. Statistical Papers, 55, 871–911.

    Article  MathSciNet  Google Scholar 

  • Drton, M., Plummer, M. (2017). A bayesian information criterion for singular models. Journal of the Royal Statistical Society Series B: Statistical Methodology, 79, 323–380.

    Article  MathSciNet  Google Scholar 

  • Forbes, F., Wraith, D. (2014). A new family of multivariate heavy-tailed distributions with variable marginal amounts of tailweight: application to robust clustering. Statistics and Computing, 24(6), 971–984.

    Article  MathSciNet  Google Scholar 

  • Fraley, C., Raftery, A. E. (2002). Model-based clustering, discriminant analysis, and density estimation. Journal of the American Statistical Association, 97(458), 611–631.

    Article  MathSciNet  Google Scholar 

  • Franczak, B. C., Browne, R. P., McNicholas, P. D. (2013). Mixtures of shifted asymmetriclaplace distributions. IEEE Transactions on Pattern Analysis and Machine Intelligence, 36, 1149–1157.

  • Fruhwirth-Schnatter, S., Celeux, G., Robert, C. P. (2019). Handbook of mixture analysis. Boca Raton: CRC Press.

    Book  Google Scholar 

  • Galimberti, G., Soffritti, G. (2014). A multivariate linear regression analysis using finite mixtures of t distributions. Computational Statistics & Data Analysis, 71, 138–150.

    Article  MathSciNet  Google Scholar 

  • Gao, X., Carroll, R. J. (2017). Data integration with high dimensionality. Biometrika, 104(2), 251–272.

    Article  MathSciNet  Google Scholar 

  • Gassiat, É. (2018). Universal Coding and Order Identification by Model Selection Methods. Cham: Springer.

    Book  Google Scholar 

  • Gassiat, E., Van Handel, R. (2012). Consistent order estimation and minimal penalties. IEEE Transactions on Information Theory, 59, 1115–1128.

    Article  MathSciNet  Google Scholar 

  • Hafidi, B., Mkhadri, A. (2010). The kullback information criterion for mixture regression models. Statistics & probability letters, 80, 807–815.

    Article  MathSciNet  Google Scholar 

  • Huang, M., Tang, S., Yao, W. (2022). Statistical inference for normal mixtures with unknown number of components. Electronic Journal of Statistics, 16, 5149–5181.

    Article  MathSciNet  Google Scholar 

  • Huang, T., Peng, H., Zhang, K. (2017). Model selection for gaussian mixture models. Statistica Sinica, 147–169.

  • Hui, F. K., Warton, D. I., Foster, S. D. (2015). Order selection in finite mixture models: complete or observed likelihood information criteria? Biometrika, 102(3), 724–730.

    Article  MathSciNet  Google Scholar 

  • Jacobs, R. A., Jordan, M. I., Nowlan, S. J., Hinton, G. E. (1991). Adaptive mixtures of local experts. Neural computation, 3, 79–87.

    Article  Google Scholar 

  • Jones, P., McLachlan, G. J. (1992). Fitting finite mixture models in a regression context. Australian Journal of Statistics, 34, 233–240.

    Article  Google Scholar 

  • Karlis, D., Xekalaki, E. (2008). The polygonal distribution. Advances in Mathematical and Statistical Modeling. New York: Springer.

  • Keribin, C. (2000). Consistent estimation of the order of mixture models. Sankhyā: The Indian Journal of Statistics. Series A, 62, 49–66.

    MathSciNet  Google Scholar 

  • Khalili, A. (2010). New estimation and feature selection methods in mixture-of-experts models. Canadian Journal of Statistics, 38, 519–539.

    Article  MathSciNet  Google Scholar 

  • Klemelä, J. (2007). Density estimation with stagewise optimization of the empirical risk. Machine learning, 67, 169–195.

    Article  Google Scholar 

  • Konishi, S., Kitagawa, G. (2008). Information criteria and statistical modeling. New York: Springer.

    Book  Google Scholar 

  • Li, J., Barron, A. (1999). Mixture density estimation. Advances in neural information processing systems

  • Lin, T. I., Lee, J. C., Hsieh, W. J. (2007). Robust mixture modeling using the skew t distribution. Statistics and computing, 17, 81–92.

  • Lindsay, B. G. (1995). Mixture models: theory, geometry, and applications. Hayward, CA: Institute of Mathematical Statistics.

    Google Scholar 

  • Lo, K., Gottardo, R. (2012). Flexible mixture modeling via the multivariate t distribution with the box-cox transformation: an alternative to the skew-t distribution. Statistics and computing, 22, 33–52.

    Article  MathSciNet  Google Scholar 

  • Lohmeyer, J., Palm, F., Reuvers, H., Urbain, J.-P. (2019). Focused information criterion for locally misspecified vector autoregressive models. Econometric Reviews, 38(7), 763–792.

    Article  MathSciNet  Google Scholar 

  • Loos, C. (2016). Analysis of Single-Cell Data: ODE Constrained Mixture Modeling and Approximate Bayesian Computation. Cham: Springer.

    Book  Google Scholar 

  • Manole, T., Ho, N. (2022). Refined convergence rates for maximum likelihood estimation under finite mixture models. International Conference on Machine Learning, 14979–15006.

  • Massart, P. (2007). Concentration Inequalities and Model Selection: École d’Été de Probabilités de Saint-Flour XXXIII - 2003 Lecture Notes in Mathematics (Vol. 1896). Berlin, Heidelberg: Springer.

  • Maugis, C., Michel, B. (2011). A non asymptotic penalized criterion for gaussian mixture model selection. ESAIM: Probability and Statistics, 15, 41–68.

    Article  MathSciNet  Google Scholar 

  • Maugis-Rabusseau, C., Michel, B. (2013). Adaptive density estimation for clustering with Gaussian mixtures. ESAIM: Probability and Statistics, 17, 698–724.

    Article  MathSciNet  Google Scholar 

  • McLachlan, G. J. (1987). On bootstrapping the likelihood ratio test statistic for the number of components in a normal mixture. Journal of the Royal Statistical Society: Series C (Applied Statistics), 36, 318–324.

    Google Scholar 

  • McLachlan, G. J., Rathnayake, S. (2014). On the number of components in a gaussian mixture model. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 4, 341–355.

    Google Scholar 

  • Mitianoudis, N., Stathaki, T. (2007). Batch and online underdetermined source separation using laplacian mixture models. IEEE Transactions on Audio, Speech, and Language Processing, 15, 1818–1832.

    Article  Google Scholar 

  • Naik, P. A., Shi, P., Tsai, C.-L. (2007). Extending the akaike information criterion to mixture regression models. Journal of the American Statistical Association, 102, 244–254.

    Article  MathSciNet  Google Scholar 

  • Ng, S. K., Xiang, L., Yau, K. K. W. (2019). Mixture Modelling for Medical and Health Sciences. Boca Raton: CRC Press.

    Book  Google Scholar 

  • Nguyen, H., Nguyen, T., Ho, N. (2023). Demystifying softmax gating function in gaussian mixture of experts. Advances in Neural Information Processing Systems, 36, 4624–4652.

    Article  Google Scholar 

  • Nguyen, H., Akbarian, P., Nguyen, T., Ho, N. (2024a). A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts. In R. Salakhutdinov, Z. Kolter, K. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp (Eds.), Proceedings of the 41st International Conference on Machine Learning,volume 235 of Proceedings of Machine Learning Research (pp. 37617–37648).: PMLR.

  • Nguyen, H., Nguyen, T., Nguyen, K., Ho, N. (2024b). Towards Convergence Rates for Parameter Estimation in Gaussian-gated Mixture of Experts. In S. Dasgupta, S. Mandt, and Y. Li (Eds.), Proceedings of The 27th International Conference on Artificial Intelligence and Statistics, volume 238 of Proceedings of Machine Learning Research (pp. 2683–2691).: PMLR.

  • Nguyen, H. D. (2024). Panic: Consistent information criteria for general model selection problems. Australian & New Zealand Journal of Statistics, 66, 441–466.

    Article  MathSciNet  Google Scholar 

  • Nguyen, H. D., Chamroukhi, F. (2018). Practical and theoretical aspects of mixture-of-experts modeling: An overview. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 8, e1246.

    Google Scholar 

  • Nguyen, H. D., McLachlan, G. J. (2016). Maximum likelihood estimation of triangular and polygonal distributions. Computational Statistics & Data Analysis, 102, 23–36.

    Article  MathSciNet  Google Scholar 

  • Nguyen, H. D., Fryer, D., McLachlan, G. J. (2023b). Order selection with confidence for finite mixture models. Journal of the Korean Statistical Society, 52, 154–184.

  • Nguyen, T., Nguyen, H. D., Chamroukhi, F., Forbes, F. (2022). A non-asymptotic approach for model selection via penalization in high-dimensional mixture of experts models. Electronic Journal of Statistics, 16(2), 4742–4822.

    Article  MathSciNet  Google Scholar 

  • Nguyen, T., Nguyen, D. N., Nguyen, H. D., Chamroukhi, F. (2023c). A non-asymptotic risk bound for model selection in a high-dimensional mixture of experts via joint rank and variable selection. Australasian Joint Conference on Artificial Intelligence, 234–245.

  • Nguyen, T., Forbes, F., Arbel, J., Duy Nguyen, H. (2024c). Bayesian nonparametric mixture of experts for inverse problems. Journal of Nonparametric Statistics, 1–60.

  • Nishii, R. (1988). Maximum likelihood principle and model selection when the true model is unspecified. Journal of Multivariate Analysis 27, 392–403.

  • Pandhare, S. C., Ramanathan, T. V. (2020a) The focussed information criterion for generalised linear regression models for time series. Australian & New Zealand Journal of Statistics, 62(4), 485–507.

  • Pandhare, S. C., Ramanathan, T. V. (2020b). The robust focused information criterion for strong mixing stochastic processes with L2-differentiable parametric densities. Statistical Inference for Stochastic Processes, 23(3), 637–663.

  • Patilea, V. (2001). Convex models, MLE and misspecification. The Annals of Statistics, 29(1), 94–123.

    Article  MathSciNet  Google Scholar 

  • Peel, D., McLachlan, G. (2000a). Finite mixture models. New York: Wiley.

  • Peel, D., McLachlan, G. J. (2000b). Robust mixture modelling using the t distribution. Statistics and computing, 10, 339–348.

  • Polymenis, A., Titterington, D. (1998). On the determination of the number of components in a mixture. Statistics & probability letters, 38, 295–298.

    Article  Google Scholar 

  • Quandt, R. E. (1972). A new approach to estimating switching regressions. Journal of the American Statistical Association, 67, 306–310.

    Article  Google Scholar 

  • Rabbani, H., Vafadust, M. (2008). Image/video denoising based on a mixture of laplace distributions with local parameters in multidimensional complex wavelet domain. Signal Processing, 88, 158–173.

    Article  Google Scholar 

  • Rakhlin, A., Panchenko, D., Mukherjee, S. (2005). Risk bounds for mixture density estimation. ESAIM: Probability and Statistics, 9, 220–229.

    Article  MathSciNet  Google Scholar 

  • Ritter, G. (2014). Robust cluster analysis and variable selection. Boca Raton: CRC Press.

    Book  Google Scholar 

  • Rockafellar, R. T., Wets, R.J.-B. (2009). Variational analysis. Berlin, Heidelberg: Springer.

    Google Scholar 

  • Schwarz, G. (1978). Estimating the dimension of a model. The Annals of Statistics, 461–464.

  • Sin, C.-Y., White, H. (1996). Information criteria for selecting possibly misspecified parametric models. Journal of Econometrics, 71, 207–225.

    Article  MathSciNet  Google Scholar 

  • Song, W., Yao, W., Xing, Y. (2014). Robust mixture regression model fitting by laplace distribution. Computational Statistics & Data Analysis, 71, 128–137.

  • Thai, T., Nguyen, T., Do, D., Ho, N., Drovandi, C. (2025). Model Selection for Gaussian-gated Gaussian Mixture of Experts Using Dendrograms of Mixing Measures. arXiv preprint arXiv:2505.13052.

  • Titterington, D. M., Smith, A. F., Makov, U. E. (1985). Statistical analysis of finite mixture distributions. Chichester: Wiley.

    Google Scholar 

  • van de Geer, S. (2000). Empirical Processes in M-estimation (Vol. 6). Cambridge: Cambridge University Press.

    Google Scholar 

  • van der Vaart, A. W. (2000). Asymptotic statistics. Cambridge: Cambridge University Press.

    Google Scholar 

  • van der Vaart, A. W., Wellner, J. A. (2023). Weak Convergence and Empirical Processes: With Applications to Statistics. Cham: Springer.

    Book  Google Scholar 

  • Vuong, Q. H. (1989). Likelihood ratio tests for model selection and non-nested hypotheses. Econometrica: Journal of the Econometric Society, 57, 307–333.

    Article  MathSciNet  Google Scholar 

  • Wainwright, M. J. (2019). High-Dimensional Statistics: A Non-Asymptotic Viewpoint. Cambridge Series in Statistical and Probabilistic Mathematics, 48. Cambridge: Cambridge University Press.

  • Wasserman, L., Ramdas, A., Balakrishnan, S. (2020). Universal inference. Proceedings of the National Academy of Sciences, 117, 16880–16890.

    Article  MathSciNet  Google Scholar 

  • Watanabe, S. (2013). A widely applicable bayesian information criterion. The Journal of Machine Learning Research, 14, 867–897.

    MathSciNet  Google Scholar 

  • Wedel, M., DeSarbo, W. S. (1995). A mixture likelihood approach for generalized linear models. Journal of classification, 12, 21–55.

    Article  Google Scholar 

  • Westerhout, J., Nguyen, T., Guo, X., Nguyen, H. D. (2024). On the asymptotic distribution of the minimum empirical risk. Forty-first International Conference on Machine Learning.

  • Yang, Y. (2005). Can the strengths of aic and bic be shared? a conflict between model identification and regression estimation. Biometrika, 92(4), 937–950.

    Article  MathSciNet  Google Scholar 

  • Yao, W., Xiang, S. (2024). Mixture Models: Parametric, Semiparametric, and New Directions. Boca Raton: CRC Press.

    Book  Google Scholar 

  • Yao, W., Wei, Y., Yu, C. (2014). Robust mixture regression using the t-distribution. Computational Statistics & Data Analysis, 71, 116–127.

  • Yuksel, S. E., Wilson, J. N., Gader, P. D. (2012). Twenty years of mixture of experts. IEEE Transactions on Neural Networks and Learning Systems, 23, 1177–1193.

    Article  Google Scholar 

Download references

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1Security firm under fire after woman is forcibly dragged from Idaho town hall: 'Violent and traumatic'0025-02-2025
2Багаевская: Ночь 18 дек, Вт0017-12-2018
3Багаевская: Утро 18 дек, Вт0017-12-2018
4Спрос на перевозки из Китая и Казахстана в Беларусь вырос в 3 раза0028-01-2025
5Изменена мера пресечения на подписку о невыезде для гендиректора ООО «ПриволжскНефтеДобыча»0028-01-2019
6Мы участвуем в проекте "Семейное дело", который запустило сообщество "Родные-Любимые. ...0022-02-2025
7«Мечел» завершил год с убытком на фоне кризиса в угольной отрасли0024-02-2025
8Затопление плавдока ПД-50 в Мурманске может повлиять на сроки ремонта кораблей0030-10-2018
9The US spends almost as much on healthcare as the rest of the world combined and has one of the worst outcomes0025-11-2022
10Сергей Цивилев проверил готовность к открытию двух детских садов в ...0009-03-2023

Классификация: Пресс-релизы. Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 8.2. Источник: link.springer.com.