Public articles linked to the same research event.
arXiv The work models dataset selection as high-dimensional regression with multiple heterogeneous sources and a weighted ridge estimator, uses only summary statistics, and provides privacy guarantees either on labels only or jointly on features and labels in terms of rho-zero-concentrated differential privacy; its main technical contribution is a deterministic equivalent of the test error that captures interactions among sample size, covariance structure, model shift, regularization, and privacy noise, allowing hyperparameters (weights and ridge regularizers) to be optimized and the usefulness of private external datasets to be judged without accessing the data itself, supported by experiments on synthetic and real-world datasets.
The work models dataset selection as high-dimensional regression with multiple heterogeneous sources and a weighted ridge estimator, uses only summary statistics, and provides privacy guarantees either on labels only or jointly on features and labels in terms of rho-zero-concentrated differential privacy; its main technical contribution is a deterministic equivalent of the test error that captures interactions among sample size, covariance structure, model shift, regularization, and privacy noise, allowing hyperparameters (weights and ridge regularizers) to be optimized and the usefulness of private external datasets to be judged without accessing the data itself, supported by experiments on synthetic and real-world datasets.
The work models dataset selection as high-dimensional regression with multiple heterogeneous sources and a weighted ridge estimator, uses only summary statistics, and provides privacy guarantees either on labels only or jointly on features and labels in terms of rho-zero-concentrated differential privacy; its main technical contribution is a deterministic equivalent of the test error that captures interactions among sample size, covariance structure, model shift, regularization, and privacy noise, allowing hyperparameters (weights and ridge regularizers) to be optimized and the usefulness of private external datasets to be judged without accessing the data itself, supported by experiments on synthetic and real-world datasets.
The work models dataset selection as high-dimensional regression with multiple heterogeneous sources and a weighted ridge estimator, uses only summary statistics, and provides privacy guarantees either on labels only or jointly on features and labels in terms of rho-zero-concentrated differential privacy; its main technical contribution is a deterministic equivalent of the test error that captures interactions among sample size, covariance structure, model shift, regularization, and privacy noise, allowing hyperparameters (weights and ridge regularizers) to be optimized and the usefulness of private external datasets to be judged without accessing the data itself, supported by experiments on synthetic and real-world datasets.