Loading…
Loading…
physics.data-an
AG-2026.07-2248
physics.data-an
S. Mitra, S. E. Lakhal, C. P. Connaughton, J. E. Sardonia, M. M. Bandi
Wind-power variability is a major challenge for the reliable integration of utility-scale wind energy into modern power systems. Although wind-speed statistics are often described by simple parametric distributions, translating these statistics into turbine-level power fluctuations is nontrivial because the relationship between wind speed and power is highly nonlinear and changes across different turbine operating regimes. Here, we develop a two-regime statistical framework for wind-power distributions. In the aerodynamic operating regime, between the cut-in and rated speeds, the turbine power follows an approximate cubic dependence on wind speed. Starting from a physically motivated Rician model for the wind-speed magnitude, we derive an analytical expression for the corresponding wind-power distribution using a nonlinear change of variables. In the control-dominated near-rated regime, where active blade-pitch and generator control regulate the turbine output, the aerodynamic transformation is no longer applicable. Instead, we characterize the power deficit relative to the rated power and show empirically that its continuous tail is well described by a bounded stretched-exponential distribution for both individual turbines and wind-farm ensembles. These results provide a physically interpretable statistical description of wind-power fluctuations across the full operational range of utility-scale wind turbines.
28 Jul 2026
1mo ago
AG-2026.07-2249
physics.data-an
S. Mitra, S. E. Lakhal, C. P. Connaughton, J. E. Sardonia, M. M. Bandi
The statistics of atmospheric wind variations are commonly modeled using Gaussian or Weibull forms, which often trade physical interpretability against statistical accuracy, especially in the distribution tails. Here we derive a Rician distribution for wind speed from a simple physical model based on two orthogonal Gaussian velocity components with a non-zero mean in the preferred direction. Using wind-speed records from four geographically distinct wind farms, we show that the Rician model consistently outperforms the Gaussian model and remains competitive with the Weibull model. The same behavior persists when the data are partitioned into monthly windows, where the Rician parameters also provide a transparent description of seasonal and geographic variability, compared to Weibull parameters. In addition, the model naturally connects Gaussian-like and Weibull-like regimes through the Rician parameter ratio $μ/σ$, making the Rician distribution a compact and physically interpretable two-parameter model for wind-speed statistics.
28 Jul 2026
1mo ago
AG-2026.06-2010
physics.data-an
Rikab Gambhir, Luisa Lucie-Smith, Jesse Thaler
We review the concepts of interpretability and explainability as they apply to machine learning in physics. We define interpretability as concerning the structural transparency of a model (the ability to understand or approximate its inner workings) and explainability as concerning the scientific content of a model (the ability to map it onto domain knowledge). We discuss the trade-offs each entails (interpretability vs. expressivity; explainability vs. adaptability), the contexts in which each is needed, and the intrinsic and post-hoc tools available for achieving them. Throughout, we emphasize that machine-learned models are subject to the same scientific questions as classical models, differing only in scale, and that interpretability and explainability are best understood as deliberate modeling choices rather than inherent properties. We also emphasize the importance of task specification and intervention plans as a core aspect of model design.
24 Jun 2026
2mo ago
AG-2026.03-1562
physics.data-an
Taikan Suehara, Takahiro Kawahara, Tomohiko Tanabe, Risako Tagami
We study the performance of the Particle Transformer (ParT) for jet flavor tagging using ILD full simulation events (1M jets) as well as fast simulation samples (10M and 1M jets). We perform 3-category ($b/c/d$), 6-category ($b/c/d/u/s/g$), and 11-category trainings (including quark--antiquark separation), incorporating multivariate hadron particle identification information from $dE/dx$ and time-of-flight. For $b$/$c$ tagging, we observe a factor of 5--10 improvement over previous BDT-based taggers, and we obtain reasonable performance for strange tagging and quark/antiquark separation. The 10M-jet fast simulation study indicates that further gains are possible with higher training statistics.
19 Mar 2026
AG-2025.12-990
physics.data-an
Javier M. Duarte, Uros Seljak, Kazu Terao
This chapter gives an overview of the core concepts of machine learning (ML) -- the use of algorithms that learn from data, identify patterns, and make predictions or decisions without being explicitly programmed -- that are relevant to particle physics with some examples of applications to the energy, intensity, cosmic, and accelerator frontiers.
11 Dec 2025
AG-2025.10-1342
physics.data-an
Masakiyo Kitazawa, ShinIchi Esumi, Takafumi Niida, Toshihiro Nonaka
We derive analytic formulas to reconstruct particle-averaged quantities from experimental results that suffer from the efficiency loss of particle measurements. These formulas are derived under the assumption that the probabilities of observing individual particles are independent. The formulas do not agree with the conventionally used intuitive formulas.
11 Oct 2025
AG-2025.08-1245
physics.data-an
Yuval Frid, Liron Barak, Pavani Jairam, Michael Kagan, Rachel Jordan Hyneman
Background modeling is one of the most critical components in high energy physics data analyses, and for smooth backgrounds it is often performed by fitting using an analytic functional form. In this paper a novel method based on Log Gaussian Cox Processes (LGCP) is introduced to model smooth backgrounds while making minimal assumptions on the underlying shape. In LGCP, samples are assumed to be drawn from a non-homogeneous Poisson process, with an intensity function drawn from a Gaussian process. Markov Chain Monte Carlo is used for optimizing the hyper parameters and drawing the final fit for the background estimate from the posterior. Synthetic experiments comparing background modeling from functional forms and the LGCP are used to compare the different methods.
15 Aug 2025
AG-2025.08-1154
physics.data-an
Juvenal Bassa, Vidya Manian, Sudhir Malik, Arghya Chattopadhyay
Jet classification in high-energy particle physics is important for understanding fundamental interactions and probing phenomena beyond the Standard Model. Jets originate from the fragmentation and hadronization of quarks and gluons, and pose a challenge for identification due to their complex, multidimensional structure. Traditional classification methods often fall short in capturing these intricacies, necessitating advanced machine learning approaches. In this paper, we employ two neural networks simultaneously as an ensemble to tag various jet types. We convert the jet data to two-dimensional histograms instead of representing them as points in a higher-dimensional space. Specifically, this ensemble approach, hereafter referred to as Ensemble Model, is used to tag jets into classes from the JetNet dataset, corresponding to: Top Quarks, Light Quarks (up or down), and W and Z bosons. For the jet classes mentioned above, we show that the Ensemble Model can be used for both binary and multi-categorical classification. This ensemble approach learns jet features by leveraging the strengths of each constituent network achieving superior performance compared to either individual network.
9 Aug 2025
AG-2025.07-1680
physics.data-an
A. Gavrikov, A. Serafini, D. Dolzhikov, A. Garfagnini, M. Gonchar, M. Grassi, R. Brugnera, V. Cerrone, L. V. D'Auria, R. M. Guizzetti, L. Lastrucci, G. Andronico, V. Antonelli, A. Barresi, D. Basilico, M. Beretta, A. Bergnoli, M. Borghesi, A. Brigatti, R. Bruno, A. Budano, B. Caccianiga, A. Cammi, R. Caruso, D. Chiesa, C. Clementi, C. Coletta, S. Dusini, A. Fabbri, G. Felici, G. Ferrante, M. G. Giammarchi, N. Giudice, N. Guardone, F. Houria, C. Landini, I. Lippi, L. Loi, P. Lombardi, F. Mantovani, S. M. Mari, A. Martini, L. Miramonti, M. Montuschi, M. Nastasi, D. Orestano, F. Ortica, A. Paoloni, L. Pelicci, E. Percalli, F. Petrucci, E. Previtali, G. Ranucci, A. C. Re, B. Ricci, A. Romani, C. Sirignano, M. Sisti, L. Stanco, E. Stanescu Farilla, V. Strati, M. D. C Torri, C. Tuvè, C. Venettacci, G. Verde, L. Votano
Precise modeling of detector energy response is crucial for next-generation neutrino experiments which present computational challenges due to lack of analytical likelihoods. We propose a solution using neural likelihood estimation within the simulation-based inference framework. We develop two complementary neural density estimators that model likelihoods of calibration data: conditional normalizing flows and a transformer-based regressor. We adopt JUNO - a large neutrino experiment - as a case study. The energy response of JUNO depends on several parameters, all of which should be tuned, given their non-linear behavior and strong correlations in the calibration data. To this end, we integrate the modeled likelihoods with Bayesian nested sampling for parameter inference, achieving uncertainties limited only by statistics with near-zero systematic biases. The normalizing flows model enables unbinned likelihood analysis, while the transformer provides an efficient binned alternative. By providing both options, our framework offers flexibility to choose the most appropriate method for specific needs. Finally, our approach establishes a template for similar applications across experimental neutrino and broader particle physics.
31 Jul 2025
AG-2025.07-1634
physics.data-an
Kristian G. Barman, Sascha Caron, Faegheh Hasibi, Eugene Shalugin, Yoris Marcet, Johannes Otte, Henk W. de Regt, Merijn Moody
We introduce a benchmark framework developed by and for the scientific community to evaluate, monitor and steer large language model development in fundamental physics. Building on philosophical concepts of scientific understanding and creativity, we develop a scoring system in which each question is scored by an expert for its correctness, difficulty, and surprise. The questions are of three forms: (i) multiple-choice questions for conceptual understanding, (ii) analytical problems requiring mathematical derivation, and (iii) openended tasks requiring complex problem solving. Our current dataset contains diverse set of examples, including a machine learning challenge to classify high-energy physics events, such as the four top quark signal. To ensure continued relevance, we propose a living benchmark, where physicists contribute questions, for instance alongside new publications. We invite contributions via: http://www.physicsbenchmarks.org/. We hope that this benchmark will enable a targeted AI development that can make a meaningful contribution to fundamental physics research.
29 Jul 2025
AG-2025.06-1138
physics.data-an
Riccardo Finotello, Vincent Lahoche, Dine Ousmane Samary
Signal detection in high dimensions is a critical challenge in data science. While standard methods based on random matrix theory provide sharp detection thresholds for finite-rank perturbations, such as the known Baik-Ben Arous-Péché (BBP) transition, they are often insufficient for realistic data exhibiting nearly continuous (extensive-rank) signal distributions that merge with the noise bulk. In this regime, typically associated with real-world scenarios such as images for computer vision tasks, the signal does not manifest as a clear outlier but as a deformation of the spectral density's geometry. We use the functional renormalisation group (FRG) framework to probe these subtle spectral deformations. Treating the empirical spectrum as an effective field theory, we define a scale-dependent "canonical dimension" that acts as a sensitive order parameter for the spectral geometry. We show that this dimension undergoes a sharp crossover, interpreted as a "dimensional phase transition", at signal-to-noise ratios significantly lower than the standard BBP threshold. This dimensional instability is shown to correlate with a spontaneous symmetry breaking in the effective potential and a deviation of eigenvector statistics from the universal Porter-Thomas distribution, confirming the consistency of the method. Such behaviour aligns with recent theoretical results on the "extensive spike model", where signal information persists inside the noise bulk before any spectral gap opens. We validate our approach on realistic datasets, demonstrating that the FRG flow consistently detects the onset of this bulk deformation. Finally, we explore a formalisation of this methodology for analysing nearly continuous spectra, proposing a heuristic criterion for signal detection and a method to estimate the number of independent noise components based on the stability of these canonical dimensions.
30 Jun 2025
AG-2025.06-1191
physics.data-an
Jay Chan, Brandon Wang, Paolo Calafiura
Particle track reconstruction is traditionally computationally challenging due to the combinatorial nature of the tracking algorithms employed. Recent developments have focused on novel algorithms with graph neural networks (GNNs), which can improve scalability. While most of these GNN-based methods require an input graph to be constructed before performing message passing, a one-shot approach called EggNet that directly takes detector spacepoints as inputs and iteratively apply graph attention networks with an evolving graph structure has been proposed. The graphs are gradually updated to improve the edge efficiency and purity, thus providing a better model performance. In this work, we evaluate the physics and computing performance of the EggNet tracking pipeline on the full TrackML dataset. We also explore different techniques to reduce constraints on computation memory and computing time.
3 Jun 2025
AG-2025.05-1451
physics.data-an
Ibrahim Elsharkawy, Yonatan Kahn
Estimating physical parameters from data is a crucial application of machine learning (ML) in the physical sciences. However, systematic uncertainties, such as detector miscalibration, induce data distribution distortions that can erode statistical precision. In both high-energy physics (HEP) and broader ML contexts, achieving uncertainty-aware parameter estimation under these domain shifts remains an open problem. In this work, we address this challenge of uncertainty-aware parameter estimation for a broad set of tasks critical for HEP. We introduce a novel approach based on Contrastive Normalizing Flows (CNFs), which achieves top performance on the HiggsML Uncertainty Challenge dataset. Building on the insight that a binary classifier can approximate the model parameter likelihood ratio, we address the practical limitations of expressivity and the high cost of simulating high-dimensional parameter grids by embedding data and parameters in a learned CNF mapping. This mapping yields a tunable contrastive distribution that enables robust classification under shifted data distributions. Through a combination of theoretical analysis and empirical evaluations, we demonstrate that CNFs, when coupled with a classifier and established frequentist techniques, provide principled parameter estimation and uncertainty quantification through classification that is robust to data distribution distortions.
13 May 2025
AG-2025.05-939
physics.data-an
Zhaohui Wang, Xiaofeng Luo
We propose a novel centrality definition-independent method for analyzing higher-order cumulants, specifically addressing the challenge of volume fluctuations that dominate in low-energy heavy-ion collisions. This method reconstructs particle number distributions using the Edgeworth expansion, with parameters optimized via a combination of differential evolution algorithm and Bayesian inference. Its effectiveness is validated using UrQMD model simulations and benchmarked against traditional approaches, including centrality definitions based on particle multiplicity. Our results show that the proposed framework yields cumulant patterns consistent with those obtained using number of participant nucleon ($N_{\text{part}}$) based centrality observables, while eliminating the conventional reliance on centrality determination. This consistency confirms the method's ability to extract genuine physical signals, thereby paving the way for probing the intrinsic thermodynamic properties of the produced medium through event-by-event fluctuations.
6 May 2025
AG-2025.04-1485
physics.data-an
Eli Gendreau-Distler, Luc Le Pottier, Haichen Wang
We explore a generative machine learning-based approach for estimating multi-dimensional probability density functions (PDFs) in a target sample using a statistically independent but related control sample - a common challenge in particle physics data analysis. The generative model must accurately reproduce individual observable distributions while preserving the correlations between them, based on the input multidimensional distribution from the control sample. Here we present a conditional normalizing flow model (CNF) based on a chain of bijectors which learns to transform unpaired simulation events to data events. We assess the performance of the CNF model in the context of LHC Higgs to diphoton analysis, where we use the CNF model to convert a Monte Carlo diphoton sample to one that models data. We show that the CNF model can accurately model complex data distributions and correlations. We also leverage the recently popularized Modified Differential Multiplier Method (MDMM) to improve the convergence of our model and assign physical meaning to usually arbitrary loss-function parameters.
15 Apr 2025
AG-2025.04-1425
physics.data-an
Roger G. Huang, Andrew Cudd, Masaki Kawaue, Tatsuya Kikawa, Benjamin Nachman, Vinicius Mikuni, Callum Wilkinson
The choice of unfolding method for a cross-section measurement is tightly coupled to the model dependence of the efficiency correction and the overall impact of cross-section modeling uncertainties in the analysis. A key issue is the dimensionality used in unfolding, as the kinematics of all outgoing particles in an event typically affect the reconstruction performance in a neutrino detector. OmniFold is an unfolding method that iteratively reweights a simulated dataset, using machine learning to utilize arbitrarily high-dimensional information, that has previously been applied to proton-proton and proton-electron datasets. This paper demonstrates OmniFold's application to a neutrino cross-section measurement for the first time using a public T2K near detector simulated dataset, comparing its performance with traditional approaches using a mock data study.
9 Apr 2025
AG-2025.03-1459
physics.data-an
Nils Sass, Hendrik Roch, Niklas Götz, Renata Krupczak, Carl B. Rosenkvist
SPARKX is an open-source Python package developed to analyze simulation data from heavy-ion collision experiments. By offering a comprehensive suite of tools, SPARKX simplifies data analysis workflows, supports multiple formats such as OSCAR2013, and integrates seamlessly with SMASH and JETSCAPE/X-SCAPE. This paper describes SPARKX's architecture, features, and applications and demonstrates its effectiveness through detailed examples and performance benchmarks. SPARKX enhances productivity and precision in relativistic kinematics studies.
12 Mar 2025
AG-2025.01-1152
physics.data-an
Jean-Francois Arguin, Georges Azuelos, Émile Baril, Ilan Bessudo, Fannie Bilodeau, Maryna Borysova, Shikma Bressler, Samuel Calvet, Julien Donini, Etienne Dreyer, Michael Kwok Lam Chu, Eva Mayer, Ethan Meszaros, Nilotpal Kakati, Bruna Pascual Dias, Joséphine Potdevin, Amit Shkuri, Muhammad Usman
The search for resonant mass bumps in invariant-mass distributions remains a cornerstone strategy for uncovering Beyond the Standard Model (BSM) physics at the Large Hadron Collider (LHC). Traditional methods often rely on predefined functional forms and exhaustive computational and human resources, limiting the scope of tested final states and selections. This work presents BumpNet, a machine learning-based approach leveraging advanced neural network architectures to generalize and enhance the Data-Directed Paradigm (DDP) for resonance searches. Trained on a diverse dataset of smoothly-falling analytical functions and realistic simulated data, BumpNet efficiently predicts statistical significance distributions across varying histogram configurations, including those derived from LHC-like conditions. The network's performance is validated against idealized likelihood ratio-based tests, showing minimal bias and strong sensitivity in detecting mass bumps across a range of scenarios. Additionally, BumpNet's application to realistic BSM scenarios highlights its capability to identify subtle signals while managing the look-elsewhere effect. These results underscore BumpNet's potential to expand the reach of resonance searches, paving the way for more comprehensive explorations of LHC data in future analyses.
9 Jan 2025
AG-2025.01-1147
physics.data-an
Kristian G. Barman, Sascha Caron, Emily Sullivan, Henk W. de Regt, Roberto Ruiz de Austri, Mieke Boon, Michael Färber, Stefan Fröse, Faegheh Hasibi, Andreas Ipp, Rukshak Kapoor, Gregor Kasieczka, Daniel Kostić, Michael Krämer, Tobias Golling, Luis G. Lopez, Jesus Marco, Sydney Otten, Pawel Pawlowski, Pietro Vischia, Erik Weber, Christoph Weniger
This paper explores ideas and provides a potential roadmap for the development and evaluation of physics-specific large-scale AI models, which we call Large Physics Models (LPMs). These models, based on foundation models such as Large Language Models (LLMs) - trained on broad data - are tailored to address the demands of physics research. LPMs can function independently or as part of an integrated framework. This framework can incorporate specialized tools, including symbolic reasoning modules for mathematical manipulations, frameworks to analyse specific experimental and simulated data, and mechanisms for synthesizing theories and scientific literature. We begin by examining whether the physics community should actively develop and refine dedicated models, rather than relying solely on commercial LLMs. We then outline how LPMs can be realized through interdisciplinary collaboration among experts in physics, computer science, and philosophy of science. To integrate these models effectively, we identify three key pillars: Development, Evaluation, and Philosophical Reflection. Development focuses on constructing models capable of processing physics texts, mathematical formulations, and diverse physical data. Evaluation assesses accuracy and reliability by testing and benchmarking. Finally, Philosophical Reflection encompasses the analysis of broader implications of LLMs in physics, including their potential to generate new scientific understanding and what novel collaboration dynamics might arise in research. Inspired by the organizational structure of experimental collaborations in particle physics, we propose a similarly interdisciplinary and collaborative approach to building and refining Large Physics Models. This roadmap provides specific objectives, defines pathways to achieve them, and identifies challenges that must be addressed to realise physics-specific large scale AI models.
9 Jan 2025
AG-2024.09-1246
physics.data-an
Nikolaos Davis
Intermittency analysis of factorial moments is a promising method used for the detection of power-law scaling in high-energy collision data. In particular, it has been employed in the search of fluctuations characteristic of the critical point (CP) of strongly interacting matter. However, intermittency analysis has been hindered by the fact that factorial moments measurements corresponding to different scales are correlated, since the same data are conventionally used to calculate them. This invalidates many assumptions involved in fitting data sets and determining the best fit values of power-law exponents. We present a novel approach to intermittency analysis, employing the well-established statistical and data science tool of Principal Component Analysis (PCA). This technique allows for the proper handling of correlations between scales without the need for subdividing the data sets available.
21 Sept 2024
AG-2024.09-779
physics.data-an
Mohammad Hossein Namjoo
Forecasting techniques for assessing the power of future experiments to discriminate between theories or discover new laws of nature are of great interest in many areas of science. In this paper, we introduce a Bayesian forecasting method using information theory. We argue that mutual information is a suitable quantity to study in this context. Besides being Bayesian, this proposal has the advantage of not relying on the choice of fiducial parameters, describing the "true" theory (which is a priori unknown), and is applicable to any probability distribution. We demonstrate that the proposed method can be used for parameter estimation and model selection, both of which are of interest concerning future experiments. We argue that mutual information has plausible interpretation in both situations. In addition, we state a number of propositions that offer information-theoretic meaning to some of the Bayesian practices such as performing multiple experiments, combining different datasets, and marginalization.
20 Sept 2024
AG-2024.09-1227
physics.data-an
Masahiko Saito, Masahiro Morinaga, Tomoe Kishimoto, Junichi Tanaka
This paper presents a parameter scan technique for BSM signal models based on normalizing flow. Normalizing flow is a type of deep learning model that transforms a simple probability distribution into a complex probability distribution as an invertible function. By learning an invertible transformation between a complex multidimensional distribution, such as experimental data observed in collider experiments, and a multidimensional normal distribution, the normalizing flow model gains the ability to sample (or generate) pseudo experimental data from random numbers and to evaluate a log-likelihood value from multidimensional observed events. The normalizing flow model can also be extended to take multidimensional conditional variables as arguments. Thus, the normalizing flow model can be used as a generator and evaluator of pseudo experimental data conditioned by the BSM model parameters. The log-likelihood value, the output of the normalizing flow model, is a function of the conditional variables. Therefore, the model can quickly calculate gradients of the log-likelihood to the conditional variables. Following this property, it is expected that the most likely set of conditional variables that reproduce the experimental data, i.e. the optimal set of parameters for the BSM model, can be efficiently searched. This paper demonstrates this on a simple dataset and discusses its limitations and future extensions.
20 Sept 2024
AG-2024.07-1359
physics.data-an
Paolo Calafiura, Jay Chan, Loic Delabrouille, Brandon Wang
Track reconstruction is a crucial task in particle experiments and is traditionally very computationally expensive due to its combinatorial nature. Recently, graph neural networks (GNNs) have emerged as a promising approach that can improve scalability. Most of these GNN-based methods, including the edge classification (EC) and the object condensation (OC) approach, require an input graph that needs to be constructed beforehand. In this work, we consider a one-shot OC approach that reconstructs particle tracks directly from a set of hits (point cloud) by recursively applying graph attention networks with an evolving graph structure. This approach iteratively updates the graphs and can better facilitate the message passing across each graph. Preliminary studies on the TrackML dataset show better track performance compared to the methods that require a fixed input graph.
18 Jul 2024
AG-2024.06-1033
physics.data-an
Camila Pazos, Shuchin Aeron, Pierre-Hugues Beauchemin, Vincent Croft, Zhengyan Huan, Martin Klassen, Taritree Wongjirad
Correcting for detector effects in experimental data, particularly through unfolding, is critical for enabling precision measurements in high-energy physics. However, traditional unfolding methods face challenges in scalability, flexibility, and dependence on simulations. We introduce a novel approach to multidimensional object-wise unfolding using conditional Denoising Diffusion Probabilistic Models (cDDPM). Our method utilizes the cDDPM for a non-iterative, flexible posterior sampling approach, incorporating distribution moments as conditioning information, which exhibits a strong inductive bias that allows it to generalize to unseen physics processes without explicitly assuming the underlying distribution. Our results highlight the potential of this method as a step towards a "universal" unfolding tool that reduces dependence on truth-level assumptions, while enabling the unfolding of a wide range of measured distributions with improved adaptability and accuracy.
3 Jun 2024
AG-2024.05-1569
physics.data-an
Etienne Dreyer, Eilam Gross, Dmitrii Kobylianskii, Vinicius Mikuni, Benjamin Nachman, Nathalie Soybelman
Detector simulation and reconstruction are a significant computational bottleneck in particle physics. We develop Particle-flow Neural Assisted Simulations (Parnassus) to address this challenge. Our deep learning model takes as input a point cloud (particles impinging on a detector) and produces a point cloud (reconstructed particles). By combining detector simulations and reconstruction into one step, we aim to minimize resource utilization and enable fast surrogate models suitable for application both inside and outside large collaborations. We demonstrate this approach using a publicly available dataset of jets passed through the full simulation and reconstruction pipeline of the CMS experiment. We show that Parnassus accurately mimics the CMS particle flow algorithm on the (statistically) same events it was trained on and can generalize to jet momentum and type outside of the training distribution.
31 May 2024
AG-2024.05-1436
physics.data-an
Jonas Spinner, Victor Bresó, Pim de Haan, Tilman Plehn, Jesse Thaler, Johann Brehmer
Extracting scientific understanding from particle-physics experiments requires solving diverse learning problems with high precision and good data efficiency. We propose the Lorentz Geometric Algebra Transformer (L-GATr), a new multi-purpose architecture for high-energy physics. L-GATr represents high-energy data in a geometric algebra over four-dimensional space-time and is equivariant under Lorentz transformations, the symmetry group of relativistic kinematics. At the same time, the architecture is a Transformer, which makes it versatile and scalable to large systems. L-GATr is first demonstrated on regression and classification tasks from particle physics. We then construct the first Lorentz-equivariant generative model: a continuous normalizing flow based on an L-GATr network, trained with Riemannian flow matching. Across our experiments, L-GATr is on par with or outperforms strong domain-specific baselines.
23 May 2024
AG-2024.05-1213
physics.data-an
Jeremy J. H. Wilkinson, Christopher G. Lester
Hypothesis testing in high dimensional data is a notoriously difficult problem without direct access to competing models' likelihood functions. This paper argues that statistical divergences can be used to quantify the difference between the population distributions of observed data and competing models, justifying their use as the basis of a hypothesis test. We go on to point out how modern techniques for functional optimization let us estimate many divergences, without the need for population likelihood functions, using samples from two distributions alone. We use a physics-based example to show how the proposed two-sample test can be implemented in practice, and discuss the necessary steps required to mature the ideas presented into an experimental framework. The code used has been made available for others to use.
10 May 2024
AG-2024.01-2286
physics.data-an
Yuan Yuan, Adrian Lozano Duran
Predicting extreme events in chaotic systems, characterized by rare but intensely fluctuating properties, is of great importance due to their impact on the performance and reliability of a wide range of systems. Some examples include weather forecasting, traffic management, power grid operations, and financial market analysis, to name a few. Methods of increasing sophistication have been developed to forecast events in these systems. However, the boundaries that define the maximum accuracy of forecasting tools are still largely unexplored from a theoretical standpoint. Here, we address the question: What is the minimum possible error in the prediction of extreme events in complex, chaotic systems? We derive the minimum probability of error in extreme event forecasting along with its information-theoretic lower and upper bounds. These bounds are universal for a given problem, in that they hold regardless of the modeling approach for extreme event prediction: from traditional linear regressions to sophisticated neural network models. The limits in predictability are obtained from the cost-sensitive Fano's and Hellman's inequalities using the Rényi entropy. The results are also connected to Takens' embedding theorem using the information can't hurt inequality. Finally, the probability of error for a forecasting model is decomposed into three sources: uncertainty in the initial conditions, hidden variables, and suboptimal modeling assumptions. The latter allows us to assess whether prediction models are operating near their maximum theoretical performance or if further improvements are possible. The bounds are applied to the prediction of extreme events in the Rössler system and the Kolmogorov flow.
29 Jan 2024
AG-2023.12-1700
physics.data-an
Vasilis Belis, Patrick Odagiu, Thea Klæboe Årrestad
The detection of out-of-distribution data points is a common task in particle physics. It is used for monitoring complex particle detectors or for identifying rare and unexpected events that may be indicative of new phenomena or physics beyond the Standard Model. Recent advances in Machine Learning for anomaly detection have encouraged the utilization of such techniques on particle physics problems. This review article provides an overview of the state-of-the-art techniques for anomaly detection in particle physics using machine learning. We discuss the challenges associated with anomaly detection in large and complex data sets, such as those produced by high-energy particle colliders, and highlight some of the successful applications of anomaly detection in particle physics experiments.
20 Dec 2023
AG-2023.12-3006
physics.data-an
Debajyoti Sengupta, Matthew Leigh, John Andrew Raine, Samuel Klein, Tobias Golling
We introduce a new technique called Drapes to enhance the sensitivity in searches for new physics at the LHC. By training diffusion models on side-band data, we show how background templates for the signal region can be generated either directly from noise, or by partially applying the diffusion process to existing data. In the partial diffusion case, data can be drawn from side-band regions, with the inverse diffusion performed for new target conditional values, or from the signal region, preserving the distribution over the conditional property that defines the signal region. We apply this technique to the hunt for resonances using the LHCO di-jet dataset, and achieve state-of-the-art performance for background template generation using high level input features. We also show how Drapes can be applied to low level inputs with jet constituents, reducing the model dependence on the choice of input observables. Using jet constituents we can further improve sensitivity to the signal process, but observe a loss in performance where the signal significance before applying any selection is below 4$σ$.
15 Dec 2023
AG-2015.06-1562
physics.data-an
Paul A. Wiggins
We have recently proposed a new information-based approach to model selection, the Frequentist Information Criterion (FIC), that reconciles information-based and frequentist inference. The purpose of this current paper is to provide a simple example of the application of this criterion and a demonstration of the natural emergence of model complexities with both AIC-like ($N^0$) and BIC-like ($\log N$) scaling with observation number $N$. The application developed is deliberately simplified to make the analysis analytically tractable.
19 Jun 2015
AG-2015.06-671
physics.data-an
Ingo Breßler, Joachim Kohlbrecher, Andreas F. Thünemann
Small-angle X-ray and neutron scattering experiments are used in many fields of the life sciences and condensed matter research to obtain answers to questions about the shape and size of nano-sized structures, typically in the range of 1 to 100 nm. It provides good statistics for large numbers of structural units for short measurement times. With the ever-increasing quantity and quality of data acquisition, the value of appropriate tools that are able to extract valuable information is steadily increasing. SASfit has been one of the mature programs for small-angle scattering data analysis available for many years. We describe the basic data processing and analysis work-flow along with recent developments in the SASfit program package (version 0.94.6). They include (i) advanced algorithms for reduction of oversampled data sets (ii) improved confidence assessment in the optimized model parameters and (iii) a flexible plug-in system for custom user-provided models. A scattering function of a mass fractal model of branched polymers in solution is provided as an example for implementing a plug-in. The new SASfit release is available for major platforms such as Windows, Linux and Mac OS X. To facilitate documentation, it includes improved indexed user documentation as well as a web-based wiki for peer collaboration and online videos for introduction of basic usage. The usage of SASfit is illustrated by interpretation of the small-angle X-ray scattering curves of monomodal gold nanoparticles (NIST reference material 8011) and bimodal silica nanoparticles (EU reference material ERM-FD-102).
9 Jun 2015
AG-2015.06-278
physics.data-an
Olivier Cavalié, François Vernotte
The Allan variance was introduced fifty years ago for analyzing the stability of frequency standards. Beside its metrological interest, it is also an estimator of the large trends of the power spectral density (PSD) of frequency deviation. For instance, the Allan variance is able to discriminate different types of noise characterized by different power laws in the PSD. But, it was also used in other fields: accelerometry, geophysics, geodesy, astrophysics and even finances! However, it seems that up to now, it has been exclusively applied for time series analysis. We propose here to use the Allan variance onto spatial data. Interferometric synthetic aperture radar (InSAR) is used in geophysics to image ground displacements in space (over the SAR image spatial coverage) and in time thank to the regular SAR image acquisitions by dedicated satellites. The main limitation of the technique is the atmospheric disturbances that affect the radar signal while traveling from the sensor to the ground and back. In this paper, we propose to use the Allan variance for analyzing spatial data from InSAR measurements. The Allan variance was computed in XY mode as well as in radial mode for detecting different types of behavior for different space-scales, in the same way as the different types of noise versus the integration time in the classical time and frequency application. We found that radial AVAR is the more appropriate way to have an estimator insensitive to the spatial axis and we applied it on SAR data acquired over eastern Turkey for the period 2003-2011. Space AVAR allowed to well characterize noise features, classically found in InSAR such as phase decorrelation producing white noise or atmospheric delays, behaving like a random walk signal. We finally applied the space AVAR to an InSAR time series to detect when the geophysical signal, here the ground motion, emerges from the noise.
3 Jun 2015
AG-2015.06-273
physics.data-an
D. J. A. Hills, A. M. Grütter, J. J. Hudson
An activity fundamental to science is building mathematical models. These models are used to both predict the results of future experiments and gain insight into the structure of the system under study. We present an algorithm that automates the model building process in a scientifically principled way. The algorithm can take observed trajectories from a wide variety of mechanical systems and, without any other prior knowledge or tuning of parameters, predict the future evolution of the system. It does this by applying the principle of least action and searching for the simplest Lagrangian that describes the system's behaviour. By generating this Lagrangian in a human interpretable form, it also provides insight into the working of the system.
3 Jun 2015
AG-2015.05-1983
physics.data-an
Paul Quincey
The recent paper by Mohr and Phillips (arXiv:1409.2794) describes several problems relating to the treatment of angle measurement within SI, the unit hertz, and quantities that can be considered countable (rather than measureable). However, the proposals that they put forward bring new problems of their own. This paper proposes alternative suggestions that solve the problems less painfully. Specifically, clarifying the text on angle in the SI brochure; relegating the hertz to a "Non-SI unit accepted for use with the International System of Units", with specific application only for "revolutions or cycles per second"; and encouraging countable quantities to be presented as pure numbers, while requiring that a sufficient description of the quantity being counted is given in the accompanying text.
27 May 2015
AG-2015.05-1530
physics.data-an
Paul A. Wiggins, Colin H. LaMont
Change-point analysis is a flexible and computationally tractable tool for the analysis of times series data from systems that transition between discrete states and whose observables are corrupted by noise. The change-point algorithm is used to identify the time indices (change points) at which the system transitions between these discrete states. We present a unified information-based approach to testing for the existence of change points. This new approach reconciles two previously disparate approaches to Change-Point Analysis (frequentist and information-based) for testing transitions between states. The resulting method is statistically principled, parameter and prior free and widely applicable to a wide range of change-point problems.
21 May 2015
AG-2015.05-284
physics.data-an
Jean Golay, Mikhail Kanevski
The size of datasets has been increasing rapidly both in terms of number of variables and number of events. As a result, the empty space phenomenon and the curse of dimensionality complicate the extraction of useful information. But, in general, data lie on non-linear manifolds of much lower dimension than that of the spaces in which they are embedded. In many pattern recognition tasks, learning these manifolds is a key issue and it requires the knowledge of their true intrinsic dimension. This paper introduces a new estimator of intrinsic dimension based on the multipoint Morisita index. It is applied to both synthetic and real datasets of varying complexities and comparisons with other existing estimators are carried out. The proposed estimator turns out to be fairly robust to sample size and noise, unaffected by edge effects, able to handle large datasets and computationally efficient.
6 May 2015
AG-2015.05-217
physics.data-an
Jun Wei Fan
In this paper I show a matrix method to calculate the exact inverse pseudopolar grid Fourier transform, and use this transform to do noise removals in the k space of pseudopolar grids. I apply the Gaussian filter to this pseudopolar grid and find the advantages of the noise removals are very excellent by using pseudopolar grid, and finally I show the Cartesian grid denoise for comparisons. The results present the signal to noise ratio and the variance are much better when doing noise removals in the pseudopolar grid than the Cartesian grid. The noise removals of pseudopolar grid or Cartesian grid are both in the k space, and all these noises are added in the real space.
4 May 2015
AG-2015.05-279
physics.data-an
Wei Huang, Yu-jian Li, Deyong Kang, Zhi Chen
The ship-rocking is a crucial factor which affects the accuracy of the ocean-based flight vehicle measurement. Here we have analyzed four groups of ship-rocking time series in horizontal and vertical directions utilizing a Hilbert based method from statistical physics. Our method gives a way to construct an analytic signal on the two-dimensional plane from a one-dimensional time series. The analytic signal share the complete property of the original time series. From the analytic signal of a time series, we have found some information of the original time series which are often hidden from the view of the conventional methods. The analytic signals of interest usually evolve very smoothly on the complex plane. In addition, the phase of the analytic signal is usually moves linearly in time. From the auto-correlation and cross-correlation functions of the original signals as well as the instantaneous amplitudes and phase increments of the analytic signals we have found that the ship-rocking in horizontal direction drives the ship-rocking in vertical direction when the ship navigates freely. And when the ship keeps a fixed navigation direction such relation disappears. Based on these resultswe could predict certain amount of future values of the ship-rocking time series based on the current and the previous values. Our predictions are as accurate as the conventional methods from stochastic processes and provide a much wider prediction time range.
4 May 2015
AG-2015.04-2199
physics.data-an
Yaroslav Nikitenko
The directional precision of the sample mean estimator was calculated analytically for the offset exponential and normal distributions in three-dimensional space both for a finite sample and for limiting cases. It was shown that the spherical projection of the sample mean of the shifted exponential distribution has connections with modified Bessel functions and with hypergeometric functions. It was shown explicitly how the distribution of the sample mean of the exponential pdf converges near the mode to the normal distribution. Approximation formulae for the distribution of the sample mean of the shifted exponential distribution and for its directional precision and for the precision of the estimation of the direction of shift of the normal distribution were obtained.
29 Apr 2015
AG-2015.04-2215
physics.data-an
P. D. Dauncey, M. Kenzie, N. Wardle, G. J. Davies
A common problem in data analysis is that the functional form, as well as the parameter values, of the underlying model which should describe a dataset is not known a priori. In these cases some extra uncertainty must be assigned to the extracted parameters of interest due to lack of exact knowledge of the functional form of the model. A method for assigning an appropriate error is presented. The method is based on considering the choice of functional form as a discrete nuisance parameter which is profiled in an analogous way to continuous nuisance parameters. The bias and coverage of this method are shown to be good when applied to a realistic example.
28 Apr 2015
AG-2015.04-1807
physics.data-an
Fenwick C. Cooper
We introduce and demonstrate two linear inverse modelling methods for systems of stochastic ODE's with accuracy that is independent of the dimensionality (number of elements) of the state vector representing the system in question. Truncation of the state space is not required. Instead we rely on the principle that perturbations decay with distance or the fact that for many systems, the state of each data point is only determined at an instant by itself and its neighbours. We further show that all necessary calculations, as well as numerical integration of the resulting linear stochastic system, require computational time and memory proportional to the dimensionality of the state vector.
28 Apr 2015
AG-2015.04-2655
physics.data-an
E. L de Santa Helena, C. M. Nascimento, G. J. L. Gerhardt
The q-Gaussians are a class of stable distributions which are present in many scientific fields, and that behave as heavy tailed distributions for an especific range of q values. The identification of these values, which are used in the description of systems, is sometimes a hard task. In this work the identification of a q-Gaussian distribution from empirical data was done by a measure of its tail weight using robust statistics. Numerical methods were used to generate artificial data, to find out the tail weight -- medcouple, and also to adjust the curve between medcouple and the q value. We showed that the medcouple value remains unchanged when the calculation is applied to data which have long memory. A routine was made to calculate the q value and its standard deviation, when applied to empirical data. It is possible to identify a q-Gaussian by the proposed methods with higher precision than in the literature for the same data sample, or as precise as found in the literature. However, in this case, it is required a smaller sample of data. We hope that this method will be able to open new ways for identifying physical phenomena that belongs to nonextensive frameworks.
27 Apr 2015
AG-2015.04-2200
physics.data-an
A. V. Lokhov, F. V. Tkachov
We review the methods of constructing confidence intervals that account for a priori information about one-sided constraints on the parameter being estimated. We show that the so-called method of sensitivity limit yields a correct solution of the problem. Derived are the solutions for the cases of a continuous distribution with non-negative estimated parameter and a discrete distribution, specifically a Poisson process with background. For both cases, the best upper limit is constructed that accounts for the a priori information. A table is provided with the confidence intervals for the parameter of Poisson distribution that correctly accounts for the information on the known value of the background along with the software for calculating the confidence intervals for any confidence levels and magnitudes of the background (the software is freely available for download via Internet).
27 Apr 2015
AG-2015.04-1459
physics.data-an
A. D. Martin, T. C. A. Molteno
We demonstrate sequential mass inference of a suspended bag of milk powder from simulated measurements of the vertical force component at the pivot while the bag is being filled. We compare the predictions of various sequential inference methods both with and without a physics model to capture the system dynamics. We find that non-augmented and augmented-state unscented Kalman filters (UKFs) in conjunction with a physics model of a pendulum of varying mass and length provide rapid and accurate predictions of the milk powder mass as a function of time. The UKFs outperform the other method tested - a particle filter. Moreover, inference methods which incorporate a physics model outperform equivalent algorithms which do not.
22 Apr 2015
AG-2015.04-2126
physics.data-an
V. M. Malyshev
The general form of solutions for parameters of interfering Breit-Wigner resonances is found. The number of solutions is determined by the properties of roots of corresponding characteristic equation and does not exceed $2^{N-1}$, where $N$ is the number of resonances. For resonances of more complicated form, provided that their amplitudes satisfy certain conditions, for any $N\ge2$ multiple solutions also exist.
17 Apr 2015
AG-2015.04-862
physics.data-an
Thomas Kreuz, Mario Mulansky, Nebojsa Bozanic
Techniques for recording large-scale neuronal spiking activity are developing very fast. This leads to an increasing demand for algorithms capable of analyzing large amounts of experimental spike train data. One of the most crucial and demanding tasks is the identification of similarity patterns with a very high temporal resolution and across different spatial scales. To address this task, in recent years three time-resolved measures of spike train synchrony have been proposed, the ISI-distance, the SPIKE-distance, and event synchronization. The Matlab source codes for calculating and visualizing these measures have been made publicly available. However, due to the many different possible representations of the results the use of these codes is rather complicated and their application requires some basic knowledge of Matlab. Thus it became desirable to provide a more user-friendly and interactive interface. Here we address this need and present SPIKY, a graphical user interface which facilitates the application of time-resolved measures of spike train synchrony to both simulated and real data. SPIKY includes implementations of the ISI-distance, the SPIKE-distance and SPIKE-synchronization (an improved and simplified extension of event synchronization) which have been optimized with respect to computation speed and memory demand. It also comprises a spike train generator and an event detector which makes it capable of analyzing continuous data. Finally, the SPIKY package includes additional complementary programs aimed at the analysis of large numbers of datasets and the estimation of significance levels.
15 Apr 2015
AG-2015.04-2462
physics.data-an
Eugene B. Postnikov, Igor M. Sokolov
We consider the problem of linear fitting of noisy data in the case of broad (say $α$-stable) distributions of random impacts ("noise"), which can lack even the first moment. This situation, common in statistical physics of small systems, in Earth sciences, in network science or in econophysics, does not allow for application of conventional Gaussian maximum-likelihood estimators resulting in usual least-squares fits. Such fits lead to large deviations of fitted parameters from their true values due to the presence of outliers. The approaches discussed here aim onto the minimization of the width of the distribution of residua. The corresponding width of the distribution can either be defined via the interquantile distance of the corresponding distributions or via the scale parameter in its characteristic function. The methods provide the robust regression even in the case of short samples with large outliers, and are equivalent to the normal least squares fit for the Gaussian noises. Our discussion is illustrated by numerical examples.
13 Apr 2015
AG-2015.04-2398
physics.data-an
Sébastien Combrexelle, Herwig Wendt, Nicolas Dobigeon, Jean-Yves Tourneret, Steve McLaughlin, Patrice Abry
Texture characterization is a central element in many image processing applications. Multifractal analysis is a useful signal and image processing tool, yet, the accurate estimation of multifractal parameters for image texture remains a challenge. This is due in the main to the fact that current estimation procedures consist of performing linear regressions across frequency scales of the two-dimensional (2D) dyadic wavelet transform, for which only a few such scales are computable for images. The strongly non-Gaussian nature of multifractal processes, combined with their complicated dependence structure, makes it difficult to develop suitable models for parameter estimation. Here, we propose a Bayesian procedure that addresses the difficulties in the estimation of the multifractality parameter. The originality of the procedure is threefold: The construction of a generic semi-parametric statistical model for the logarithm of wavelet leaders; the formulation of Bayesian estimators that are associated with this model and the set of parameter values admitted by multifractal theory; the exploitation of a suitable Whittle approximation within the Bayesian model which enables the otherwise infeasible evaluation of the posterior distribution associated with the model. Performance is assessed numerically for several 2D multifractal processes, for several image sizes and a large range of process parameters. The procedure yields significant benefits over current benchmark estimators in terms of estimation performance and ability to discriminate between the two most commonly used classes of multifractal process models. The gains in performance are particularly pronounced for small image sizes, notably enabling for the first time the analysis of image patches as small as 64x64 pixels.
9 Apr 2015
AG-2015.03-1995
physics.data-an
Joseph R. Iafrate, Steven J. Miller, Frederick W. Strauch
A statistical model for the fragmentation of a conserved quantity is analyzed, using the principle of maximum entropy and the theory of partitions. Upper and lower bounds for the restricted partitioning problem are derived and applied to the distribution of fragments. The resulting power law directly leads to Benford's law for the first digits of the parts.
28 Mar 2015
AG-2015.03-2454
physics.data-an
S. M. Abrarov, B. M. Quine
We present a rational approximation for rapid and accurate computation of the Voigt function, obtained by residue calculus. The computational test reveals that with only $16$ summation terms this approximation provides average accuracy ${10^{- 14}}$ over a wide domain of practical interest $0 < x < 40,000$ and ${10^{- 4}} < y < {10^2}$ for applications using the HITRAN molecular spectroscopic database. The proposed rational approximation takes less than half the computation time of that required by Weideman's rational approximation. Algorithmic stability is achieved due to absence of the poles at $y \geqslant 0$ and $ - \infty < x < \infty $.
27 Mar 2015
AG-2015.03-1742
physics.data-an
Tiago P. Peixoto
The effort to understand network systems in increasing detail has resulted in a diversity of methods designed to extract their large-scale structure from data. Unfortunately, many of these methods yield diverging descriptions of the same network, making both the comparison and understanding of their results a difficult challenge. A possible solution to this outstanding issue is to shift the focus away from ad hoc methods and move towards more principled approaches based on statistical inference of generative models. As a result, we face instead the more well-defined task of selecting between competing generative processes, which can be done under a unified probabilistic framework. Here, we consider the comparison between a variety of generative models including features such as degree correction, where nodes with arbitrary degrees can belong to the same group, and community overlap, where nodes are allowed to belong to more than one group. Because such model variants possess an increasing number of parameters, they become prone to overfitting. In this work, we present a method of model selection based on the minimum description length criterion and posterior odds ratios that is capable of fully accounting for the increased degrees of freedom of the larger models, and selects the best one according to the statistical evidence available in the data. In applying this method to many empirical unweighted networks from different fields, we observe that community overlap is very often not supported by statistical evidence and is selected as a better model only for a minority of them. On the other hand, we find that degree correction tends to be almost universally favored by the available data, implying that intrinsic node proprieties (as opposed to group properties) are often an essential ingredient of network formation.
26 Mar 2015
AG-2015.03-1849
physics.data-an
Ronald F. Fox, Theodore P. Hill
A recent article by Alexopoulos and Leontsinis presented empirical evidence that the first digits of the distances to galaxies are a reasonably good fit to the probabilities predicted by Benford's law, the well known logarithmic statistical distribution of significant digits. The purpose of the present article is to give a theoretical explanation, based on Hubble's law and mathematical properties of Benford's law, why galaxy distances might be expected to follow Benford's law. The new galaxy-distance law derived here, which is robust with respect to change of scale and base, to additive and multiplicative computational or observational errors, and to variability of the Hubble constant in both time and space, predicts that conformity to Benford's law will improve as more data on distances to galaxies becomes available. Conversely, with the logical derivation of this law presented here, the recent empirical observations may be viewed as independent evidence of the validity of Hubble's law.
26 Mar 2015
AG-2015.03-1775
physics.data-an
Kyle Cranmer
This document is a pedagogical introduction to statistics for particle physics. Emphasis is placed on the terminology, concepts, and methods being used at the Large Hadron Collider. The document addresses both the statistical tests applied to a model of the data and the modeling itself.
26 Mar 2015
AG-2015.03-2178
physics.data-an
Aglae Kellerer
The structure function is a useful quantity to characterize wavefront distortions. We derive expressions for the structure functions of the averaged wavefront phase and slopes. The expressions are valid within the inertial range of atmospheric turbulence, and are meant to serve as engineering formulae when reconstructing profiles of the atmospheric turbulence, specifically in the context of atmospheric profiling instruments (e.g. SLODAR and S-DIMM+) and multi-conjugate adaptive optical systems.
26 Mar 2015
AG-2015.03-1403
physics.data-an
Ardeshir M. Ebtehaj, Rafael L. Bras, Efi Foufoula-Georgiou
For precipitation retrievals over land, using satellite measurements in microwave bands, it is important to properly discriminate the weak rainfall signals from strong and highly variable background surface emission. Traditionally, land rainfall retrieval methods often rely on a weak signal of rainfall scattering on high-frequency channels (85 GHz) and make use of empirical thresholding and regression-based techniques. Due to the increased ground surface signal interference, precipitation retrieval over radiometrically complex land surfaces, especially over snow-covered lands, deserts and coastal areas, is of particular challenge for this class of retrieval techniques. This paper evaluates the results by the recently proposed Shrunken locally linear embedding Algorithm for Retrieval of Precipitation (ShARP), over a radiometrically complex terrain and coastal areas using the data provided by the Tropical Rainfall Measuring Mission (TRMM) satellite. To this end, the ShARP retrieval experiments are performed over a region in Southeast Asia, partly covering the Tibetan Highlands, Himalayas, Ganges-Brahmaputra-Meghna river basins and its delta. We elucidate promising results by ShARP over snow covered land surfaces and at the vicinity of coastlines, in comparison with the land rainfall retrievals of the standard TRMM-2A12 product. Specifically, using the TRMM-2A25 radar product as a reference, we provide evidence that the ShARP algorithm can significantly reduce the rainfall over estimation due to the background snow contamination and markedly improve detection and retrieval of rainfall at the vicinity of coastlines. During the calendar year 2013, we demonstrate that over the study domain the root mean squared difference can be reduced up to 38% annually, while the reduction can reach up to 70% during the cold months.
19 Mar 2015
AG-2015.03-730
physics.data-an
Rüdiger Kürsten
In a recent paper [2] the author introduced and investigated a random walk model similar to a model introduced in [1]. In these models the increment of the random walk depends on the complete past of the process. In this note I will point out that the models considered in [1] and [2] can be mapped onto each other one to one. They can be defined on a common probability space and hence all expectation values of the model [2] with parameter p are equal to the ones of [1] with a corresponding parameter $\tilde{p}$.
11 Mar 2015
AG-2015.03-2245
physics.data-an
Gerrit Ansmann
We explore the mathematical foundations of the vector space of physical dimensions introduced in A. Maksymowicz, Am. J. Phys. 44, 1976, and extend this formalism to the vector space of physical values. As different unit systems correspond to different bases of this vector space, our formalism may find use for introducing the concept of natural units and transforming physical values between unit systems.
9 Mar 2015
AG-2015.03-364
physics.data-an
Sylvain Fichet
We introduce a new kind of likelihood function based on the sequence of moments of the data distribution. Both binned and unbinned data samples are discussed, and the multivariate case is also derived. Building on this approach we lay out the formalism of shape analysis for signal searches. In addition to moment-based likelihoods, standard likelihoods and approximate statistical tests are provided. Enough material is included to make the paper self-contained from the perspective of shape analysis. We argue that the moment-based likelihoods can advantageously replace unbinned standard likelihoods for the search of non-local signals, by avoiding the step of fitting Monte-Carlo generated distributions. This benefit increases with the number of variables simultaneously analyzed. The moment-based signal search is exemplified and tested in various 1D toy models mimicking typical high-energy signal--background configurations. Moment-based techniques should be particularly appropriate for the searches for effective operators at the LHC.
6 Mar 2015
AG-2015.03-028
physics.data-an
Sebastian Dorn, Torsten A. Enßlin, Maksim Greiner, Marco Selig, Vanessa Boehm
The calibration of a measurement device is crucial for every scientific experiment, where a signal has to be inferred from data. We present CURE, the calibration uncertainty renormalized estimator, to reconstruct a signal and simultaneously the instrument's calibration from the same data without knowing the exact calibration, but its covariance structure. The idea of CURE, developed in the framework of information field theory, is starting with an assumed calibration to successively include more and more portions of calibration uncertainty into the signal inference equations and to absorb the resulting corrections into renormalized signal (and calibration) solutions. Thereby, the signal inference and calibration problem turns into solving a single system of ordinary differential equations and can be identified with common resummation techniques used in field theories. We verify CURE by applying it to a simplistic toy example and compare it against existent self-calibration schemes, Wiener filter solutions, and Markov Chain Monte Carlo sampling. We conclude that the method is able to keep up in accuracy with the best self-calibration methods and serves as a non-iterative alternative to it.
2 Mar 2015
AG-2015.03-2266
physics.data-an
Diego Casadei
The objective Bayesian treatment of a model representing two independent Poisson processes, labelled as "signal" and "background" and both contributing additively to the total number of counted events, is considered. It is shown that the reference prior for the parameter of interest (the signal intensity) can be well approximated by the widely (ab)used flat prior only when the expected background is very high. On the other hand, a very simple approximation (the limiting form of the reference prior for perfect prior background knowledge) can be safely used over a large portion of the background parameters space. The resulting approximate reference posterior is a Gamma density whose parameters are related to the observed counts. This limiting form is simpler than the result obtained with a flat prior, with the additional advantage of representing a much closer approximation to the reference posterior in all cases. Hence such limiting prior should be considered a better default or conventional prior than the uniform prior. On the computing side, it is shown that a 2-parameter fitting function is able to reproduce extremely well the reference prior for any background prior. Thus, it can be useful in applications requiring the evaluation of the reference prior for a very large number of times. [The published version JINST 9 (2014) T10006 has a typo in the normalization $N$ of eq.(2.6) that is fixed here.]
2 Mar 2015
AG-2015.03-2575
physics.data-an
O. López-Corona
In this talk I explained briefly the advantages of using genetic algorithms on any measured data but specially astronomical ones. This kind of algorithms are not only a better computational paradigm, but they also allow for a more profound data treatment enhancing theoretical developments. As an example, I will use the SNIa cosmological data to fit the extended metric theories of gravity of Carranza et al. (2013, 2014) showing that the best parameters combination deviate from theoretical predicted ones by a minimal amount. This means that these kind of gravitational extensions are statistically robust and show that no dark matter and/or energy is required to explain the observations.
2 Mar 2015
AG-2015.02-2385
physics.data-an
Nikolai Gagunashvili
Weighted histograms are used for the estimation of probability density functions. Computer simulation is the main domain of application of this type of histogram. A review of chi-square goodness of fit tests for weighted histograms is presented in this paper. Improvements are proposed to these tests that have size more close to its nominal value. Numerical examples are presented in this paper for evaluation of tests and to demonstrate various applications of tests.
19 Feb 2015
AG-2015.02-1144
physics.data-an
Nico Reinke, André Fuchs, Wided Medjroubi, Pedro G. Lind, Matthias Wächter, Joachim Peinke
We describe a simple stochastic method, so-called Langevin approach, which enables one to extract evolution equations of stochastic variables from a set of measurements. Our method is parameter-free and it is based on the nonlinear Langevin equation. Moreover, it can be applied not only to processes in time, but also to processes in scale, given that the data available shows ergodicity. This chapter introduces the mathematical foundations of the Langevin approach and describes how to implement it numerically. A specific application of the method is presented, namely to a turbulent velocity field measured in the laboratory, retrieving the corresponding energy cascade and comparing with the results from a computational simulation of that experiment. In addition, we describe a physical interpretation bridging between processes in time and in scale. Finally, we describe extensions of the method for time series reconstruction and applications to other fields such as finance, medicine, geophysics and renewable energies.
18 Feb 2015
AG-2015.02-1453
physics.data-an
Thomas Markovich, Samuel M. Blau, Jacob N. Sanders, Alan Aspuru-Guzik
Signal processing techniques have been developed that use different strategies to bypass the Nyquist sampling theorem in order to recover more information than a traditional discrete Fourier transform. Here we examine three such methods: filter diagonalization, compressed sensing, and super-resolution. We apply them to a broad range of signal forms commonly found in science and engineering in order to discover when and how each method can be used most profitably. We find that filter diagonalization provides the best results for Lorentzian signals, while compressed sensing and super-resolution perform better for arbitrary signals.
17 Feb 2015
AG-2015.02-706
physics.data-an
G. D'Amico, F. Petroni, F. Prattico
Modeling wind speed is one of the key element when dealing with the production of energy through wind turbines. A good model can be used for forecasting, site evaluation, turbines design and many other purposes. In this work we are interested in the analysis of the future financial cash flows generated by selling the electrical energy produced. We apply an indexed semi-Markov model of wind speed that has been shown, in previous investigation, to reproduce accurately the statistical behavior of wind speed. The model is applied to the evaluation of financial indicators like the Internal Rate of Return, semi-Elasticity and relative Convexity that are widely used for the assessment of the profitability of an investment and for the measurement and analysis of interest rate risk. We compare the computation of these indicators for real and synthetic data. Moreover, we propose a new indicator that can be used to compare the degree of utilization of different power plants.
11 Feb 2015
AG-2015.02-892
physics.data-an
Keijo Hämäläinen, Lauri Harhanen, Aki Kallonen, Antti Kujanpää, Esa Niemi, Samuli Siltanen
This is the documentation of the tomographic X-ray data of a walnut made available at http://www.fips.fi/dataset.php . The data can be freely used for scientific purposes with appropriate references to the data and to this document in arXiv. The data set consists of (1) the X-ray sinogram of a single 2D slice of the walnut with three different resolutions and (2) the corresponding measurement matrices modeling the linear operation of the X-ray transform. Each of these sinograms was obtained from a measured 120-projection fan-beam sinogram by down-sampling and taking logarithms. The original (measured) sinogram is also provided in its original form and resolution. In addition, a larger set of 1200 projections of the same walnut was measured and a high-resolution filtered back-projection reconstruction was computed from this data; both the sinogram and the FBP reconstruction are included in the data set, the latter serving as a ground truth reconstruction.
11 Feb 2015
AG-2015.02-2051
physics.data-an
Luca M. Ghiringhelli, Jan Vybiral, Sergey V. Levchenko, Claudia Draxl, Matthias Scheffler
Statistical learning of materials properties or functions so far starts with a largely silent, non-challenged step: the choice of the set of descriptive parameters (termed descriptor). However, when the scientific connection between the descriptor and the actuating mechanisms is unclear, causality of the learned descriptor-property relation is uncertain. Thus, trustful prediction of new promising materials, identification of anomalies, and scientific advancement are doubtful. We analyse this issue and define requirements for a suited descriptor. For a classical example, the energy difference of zincblende/wurtzite and rocksalt semiconductors, we demonstrate how a meaningful descriptor can be found systematically.
5 Feb 2015
AG-2015.02-109
physics.data-an
Steven D. Biller, Scott M. Oser
The behaviors of various confidence/credible interval constructions are explored, particularly in the region of low statistics where methods diverge most. We highlight a number of challenges, such as the treatment of nuisance parameters, and common misconceptions associated with such constructions. An informal survey of the literature suggests that confidence intervals are not always defined in relevant ways and are too often misinterpreted and/or misapplied. This can lead to seemingly paradoxical behaviours and flawed comparisons regarding the relevance of experimental results. We therefore conclude that there is a need for a more pragmatic strategy which recognizes that, while it is critical to objectively convey the information content of the data, there is also a strong desire to derive bounds on models and a natural instinct to interpret things this way. Accordingly, we attempt to put aside philosophical biases in favor of a practical view to propose a more transparent and self-consistent approach that better addresses these issues.
3 Feb 2015
AG-2015.01-1728
physics.data-an
Vincent Wens
Network theory and inverse modeling are two standard tools of applied physics, whose combination is needed when studying the dynamical organization of spatially distributed systems from indirect measurements. However, the associated connectivity estimation may be affected by spatial leakage, an artifact of inverse modeling that limits the interpretability of network analysis. This paper investigates general analytical aspects pertaining to this issue. First, the existence of spatial leakage is derived from the topological structure of inverse operators. Then, the geometry of spatial leakage is modeled and used to define a geometric correction scheme, which limits spatial leakage effects in connectivity estimation. Finally, this new approach for network analysis is compared analytically to existing methods based on linear regressions, which are shown to yield biased coupling estimates.
29 Jan 2015
AG-2015.01-1199
physics.data-an
Kevin H. Knuth
High temporal resolution measurements of human brain activity can be performed by recording the electric potentials on the scalp surface (electroencephalography, EEG), or by recording the magnetic fields near the surface of the head (magnetoencephalography, MEG). The analysis of the data is problematic due to the fact that multiple neural generators may be simultaneously active and the potentials and magnetic fields from these sources are superimposed on the detectors. It is highly desirable to un-mix the data into signals representing the behaviors of the original individual generators. This general problem is called blind source separation and several recent techniques utilizing maximum entropy, minimum mutual information, and maximum likelihood estimation have been applied. These techniques have had much success in separating signals such as natural sounds or speech, but appear to be ineffective when applied to EEG or MEG signals. Many of these techniques implicitly assume that the source distributions have a large kurtosis, whereas an analysis of EEG/MEG signals reveals that the distributions are multimodal. This suggests that more effective separation techniques could be designed for EEG and MEG signals.
21 Jan 2015
AG-2015.01-1242
physics.data-an
Alexander Heifetz, Sasan Bakhtiari, Apostolos C. Raptis
We have investigated the application of Compressive Sensing (CS) computational method to simultaneous compression and encryption of gamma spectra measured with NaI(Tl) detector during wide area search survey applications. Our numerical experiments have demonstrated secure encryption and nearly lossless recovery of gamma spectra coded and decoded with CS routines.
21 Jan 2015
AG-2015.01-1200
physics.data-an
Kevin H. Knuth, Herbert G. Vaughan
We consider two areas of research that have been developing in parallel over the last decade: blind source separation (BSS) and electromagnetic source estimation (ESE). BSS deals with the recovery of source signals when only mixtures of signals can be obtained from an array of detectors and the only prior knowledge consists of some information about the nature of the source signals. On the other hand, ESE utilizes knowledge of the electromagnetic forward problem to assign source signals to their respective generators, while information about the signals themselves is typically ignored. We demonstrate that these two techniques can be derived from the same starting point using the Bayesian formalism. This suggests a means by which new algorithms can be developed that utilize as much relevant information as possible. We also briefly mention some preliminary work that supports the value of integrating information used by these two techniques and review the kinds of information that may be useful in addressing the ESE problem.
21 Jan 2015
AG-2015.01-1219
physics.data-an
Yui Shiozawa, Bruce N. Miller, Jean-Louis Rouet
While the numerical methods which utilizes partitions of equal-size, including the box-counting method, remain the most popular choice for computing the generalized dimension of multifractal sets, two mass- oriented methods are investigated by applying them to the one-dimensional generalized Cantor set. We show that both mass-oriented methods generate relatively good results for generalized dimensions for important cases where the box-counting method is known to fail. Both the strengths and limitations of the methods are also discussed.
21 Jan 2015
AG-2015.01-2869
physics.data-an
Francois Vernotte, Eric Lantz
1/f noise is very common but is difficult to handle in a metrological way. After having recalled the main characteristics of stongly correlated noise, this paper will determine relationships giving confidence intervals over the arithmetic mean and the linear drift parameters. A complete example of processing of an actual measurement sequence affected by 1/f noise will be given.
19 Jan 2015
AG-2015.01-878
physics.data-an
Luca Lista
The best linear unbiased estimator (BLUE) is a popular statistical method adopted to combine multiple measurements of the same observable taking into account individual uncertainties and their correlation. The method is unbiased by construction if the true uncertainties and their correlation are known, but it may exhibit a bias if uncertainty estimates are used in place of the true ones, in particular if those estimated uncertainties depend on measured values. This is the case for instance when contributions to the total uncertainty are known as relative uncertainties. In those cases, an iterative application of the BLUE method may reduce the bias of the combined measurement. The impact of the iterative approach compared to the standard BLUE application is studied for a wide range of possible values of uncertainties and their correlation in the case of the combination of two measurements.
16 Jan 2015
AG-2015.01-539
physics.data-an
Igor Volobouev
Error propagation formulae are derived for the expectation-maximization iterative unfolding algorithm regularized by a smoothing step. The effective number of parameters in the fit to the observed data is defined for unfolding procedures. Based upon this definition, the Akaike information criterion is proposed as a principle for choosing the smoothing parameters in an automatic, data-dependent manner. The performance and the frequentist coverage of the resulting method are investigated using simulated samples. A number of issues of general relevance to all unfolding techniques are discussed, including irreducible bias, uncertainty increase due to a data-dependent choice of regularization strength, and presentation of results.
9 Jan 2015
AG-2015.01-241
physics.data-an
J. Li, P. Vignal, S. Sun, V. M. Calo
In Markov Chain Monte Carlo (MCMC) simulations, the thermal equilibria quantities are estimated by ensemble average over a sample set containing a large number of correlated samples. These samples are selected in accordance with the probability distribution function, known from the partition function of equilibrium state. As the stochastic error of the simulation results is significant, it is desirable to understand the variance of the estimation by ensemble average, which depends on the sample size (i.e., the total number of samples in the set) and the sampling interval (i.e., cycle number between two consecutive samples). Although large sample sizes reduce the variance, they increase the computational cost of the simulation. For a given CPU time, the sample size can be reduced greatly by increasing the sampling interval, while having the corresponding increase in variance be negligible if the original sampling interval is very small. In this work, we report a few general rules that relate the variance with the sample size and the sampling interval. These results are observed and confirmed numerically. These variance rules are derived for the MCMC method but are also valid for the correlated samples obtained using other Monte Carlo methods. The main contribution of this work includes the theoretical proof of these numerical observations and the set of assumptions that lead to them.
7 Jan 2015
AG-2015.01-2196
physics.data-an
Nikolai Gagunashvili
A procedure based on a Mixture Density Model for correcting experimental data for distortions due to finite resolution and limited detector acceptance is presented. Addressing the case that the solution is known to be non-negative, in the approach presented here, the true distribution is estimated by a weighted sum of probability density functions with positive weights and with the width of the densities acting as a regularisation parameter responsible for the smoothness of the result. To obtain better smoothing in less populated regions, the width parameter is chosen inversely proportional to the square root of the estimated density. Furthermore, the non-negative garrotte method is used to find the most economic representation of the solution. Cross-validation is employed to determine the optimal values of the resolution and garrotte parameters. The proposed approach is directly applicable to multidimensional problems. Numerical examples in one and two dimensions are presented to illustrate the procedure.
5 Jan 2015
AG-2014.12-2613
physics.data-an
Ardeshir Mohammad Ebtehaj, Efi Foufoula-Georgiou, Gilad Lerman, Rafael Luis Bras
We demonstrate that the global fields of temperature, humidity and geopotential heights admit a nearly sparse representation in the wavelet domain, offering a viable path forward to explore new paradigms of sparsity-promoting data assimilation and compressive recovery of land surface-atmospheric states from space. We illustrate this idea using retrieval products of the Atmospheric Infrared Sounder (AIRS) and Advanced Microwave Sounding Unit (AMSU) on board the Aqua satellite. The results reveal that the sparsity of the fields of temperature is relatively pressure-independent while atmospheric humidity and geopotential heights are typically sparser at lower and higher pressure levels, respectively. We provide evidence that these land-atmospheric states can be accurately estimated using a small set of measurements by taking advantage of their sparsity prior.
30 Dec 2014
AG-2014.12-1839
physics.data-an
Alessandro Petrolini
The least squares fit to a straight line, when both variables are affected by all equal uncorrelated errors, leads to very simple results for both the estimated parameters and their standard errors, of widespread applicability. In this paper several formulas are derived, presenting a full set of results about the estimated parameters and their standard errors. All the results have been validated with extensive Monte Carlo simulations. The emphasis of the paper is on the calculation and properties of the best-fit parameters and their standard errors.
27 Dec 2014
AG-2014.12-1921
physics.data-an
Marcos W. S. Oliveira, Dalcimar Casanova, João B. Florindo, Odemir Martinez Bruno
This work proposes to obtain novel fractal descriptors from gray-level texture images by combining information from interior and boundary measures of the Minkowski dilation applied to the texture surface. At first, the image is converted into a surface where the height of each point is the gray intensity of the respective pixel in that position in the image. Thus, this surface is morphologically dilated by spheres. The radius of such spheres is ranged within an interval and the volume and the external area of the dilated structure are computed for each radius. The final descriptors are given by such measures concatenated and subject to a canonical transform to reduce the dimensionality. The proposal is an enhancement to the classical Bouligand-Minkowski fractal descriptors, where only the volume (interior) information is considered. As different structures may have the same volume, but not the same area, the proposal yields to more rich descriptors as confirmed by results on the classification of benchmark databases.
26 Dec 2014
AG-2014.12-1619
physics.data-an
J. P. Huang
If a person looks at WHITE paper through BLUE glasses, the paper will become BLUE in the eye of the person. Likewise, in the current study of big data which play the same role as the white paper being looked at, various statistical methods just serve as the blue glasses. That is, results obtained from big data often depend on the statistical methods in use, which may often defy reality. Here I suggest using physical ideas and methods to overcome this problem to the greatest extent. This suggestion is helpful to development and application of big data.
22 Dec 2014
AG-2014.12-1594
physics.data-an
A. Z. Gorski, M. Stroz, P. Oswiecimka, J. Skrzat
The box-counting (BC) algorithm is applied to calculate fractal dimensions of four fractal sets. The sets are contaminated with an additive noise with amplitude $γ= 10^{-5} ÷10^{-1}$. The accuracy of calculated numerical values of the fractal dimensions is analyzed as a function of $γ$ for different sizes of the data sample ($n_{tot}$). In particular, it has been found that a tiny ($10^{-5}$) addition of noise generates much larger (three orders of magnitude) error of the calculated fractal exponents. Natural saturation of the error for larger noise values prohibits the power-like scaling. Moreover, the noise effect cannot be cured by taking larger data samples.
20 Dec 2014
AG-2014.12-1419
physics.data-an
Federico Colecchia
Background properties in experimental particle physics are typically estimated using large data sets. However, different events can exhibit different features because of the quantum mechanical nature of the underlying physics processes. While signal and background fractions in a given data set can be evaluated using a maximum likelihood estimator, the shapes of the corresponding distributions are traditionally obtained using high-statistics control samples, which normally neglects the effect of fluctuations. On the other hand, if it was possible to subtract background using templates that take fluctuations into account, this would be expected to improve the resolution of the observables of interest, and to reduce systematics depending on the analysis. This study is an initial step in this direction. We propose a novel algorithm inspired by the Gibbs sampler that makes it possible to estimate the shapes of signal and background probability density functions from a given collection of particles, using control sample templates as initial conditions and refining them to take into account the effect of fluctuations. Results on Monte Carlo data are presented, and the prospects for future development are discussed.
19 Dec 2014
AG-2014.12-1424
physics.data-an
Federico Colecchia
Low-energy strong interactions are a major source of background at hadron colliders, and methods of subtracting the associated energy flow are well established in the field. Traditional approaches treat the contamination as diffuse, and estimate background energy levels either by averaging over large data sets or by restricting to given kinematic regions inside individual collision events. On the other hand, more recent techniques take into account the discrete nature of background, most notably by exploiting the presence of substructure inside hard jets, i.e. inside collections of particles originating from scattered hard quarks and gluons. However, none of the existing methods subtract background at the level of individual particles inside events. We illustrate the use of an algorithm that can enable particle-by-particle background discrimination at the Large Hadron Collider, and we envisage this as the basis for a novel event filtering procedure upstream of the official jet reconstruction pipelines. Our hope is that this new technique will improve physics analysis when used in combination with state-of-the-art algorithms in high-luminosity hadron collider environments.
19 Dec 2014
AG-2014.12-1475
physics.data-an
Kevin H. Knuth, Deniz Gençağa, William B. Rossow
Information-theoretic quantities, such as entropy, are used to quantify the amount of information a given variable provides. Entropies can be used together to compute the mutual information, which quantifies the amount of information two variables share. However, accurately estimating these quantities from data is extremely challenging. We have developed a set of computational techniques that allow one to accurately compute marginal and joint entropies. These algorithms are probabilistic in nature and thus provide information on the uncertainty in our estimates, which enable us to establish statistical significance of our findings. We demonstrate these methods by identifying relations between cloud data from the International Satellite Cloud Climatology Project (ISCCP) and data from other sources, such as equatorial pacific sea surface temperatures (SST).
19 Dec 2014
AG-2014.12-1499
physics.data-an
Sergio Callegari
Spread-spectrum signals are increasingly adopted in fields including communications, testing of electronic systems, Electro-Magnetic Compatibility (EMC) enhancement, ultrasonic non-destructive testing. This paper considers the synthesis of constant-envelope band-pass wave-forms with preassigned spectra via an FM technique using only a limited number of frequencies. In particular, an optimization-based approach for the selection of appropriate modulation parameters and statistical features of the modulating waveform is proposed. By example, it is shown that the design problem generally admits multiple local optima, but can still be managed with relative ease since the local optima can typically be scanned by changing the initial setting of a single parameter.
19 Dec 2014
AG-2014.12-3864
physics.data-an
Ariel Caticha
To what extent can we distinguish one probability distribution from another? Are there quantitative measures of distinguishability? The goal of this tutorial is to approach such questions by introducing the notion of the "distance" between two probability distributions and exploring some basic ideas of such an "information geometry".
17 Dec 2014
AG-2014.12-2323
physics.data-an
Federico Colecchia
When the number of events associated with a signal process is estimated in particle physics, it is common practice to extrapolate background distributions from control regions to a predefined signal window. This allows accurate estimation of the expected, or average, number of background events under the signal. However, in general, the actual number of background events can deviate from the average due to fluctuations in the data. Such a difference can be sizable when compared to the number of signal events in the early stages of data analysis following the observation of a new particle, as well as in the analysis of rare decay channels. We report on the development of a data-driven technique that aims to estimate the actual, as opposed to the expected, number of background events in a predefined signal window. We discuss results on toy Monte Carlo data and provide a preliminary estimate of systematic uncertainty.
16 Dec 2014
AG-2014.12-496
physics.data-an
Bernd Lehle, Joachim Peinke
The stochastic properties of a Langevin-type Markov process can be extracted from a given time series by a Markov analysis. Also processes that obey a stochastically forced second order differential equation can be analyzed this way by employing a particular embedding approach: To obtain a Markovian process in 2N dimensions from a non Markovian signal in N dimensions, the system is described in a phase space that is extended by the temporal derivative of the signal. For a discrete time series, however, this derivative can only be calculated by a differencing scheme, which introduces an error. If the effects of this error are not accounted for, this leads to systematic errors in the estimation of the drift- and diffusion functions of the process. In this paper we will analyze these errors and we will propose an approach that correctly accounts for them. This approach allows an accurate parameter estimation and, additionally, is able to cope with weak measurement noise, which may be superimposed to a given time series.
6 Dec 2014
AG-2014.12-396
physics.data-an
Ingo Breßler, Brian R. Pauw, Andreas Thünemann
A reliable and user-friendly characterisation of nano-objects in a target material is presented here in the form of a software data analysis package for interpreting small-angle X-ray scattering (SAXS) patterns. When provided with data on absolute scale with reasonable uncertainty estimates, the software outputs (size) distributions in absolute volume fractions complete with uncertainty estimates and minimum evidence limits, and outputs all distribution modes of a user definable range of one or more model parameters. A multitude of models are included, including prolate and oblate nanoparticles, core-shell objects, polymer models (Gaussian chain and Kholodenko worm) and a model for densely packed spheres (using the LMA-PY approximations). The McSAS software can furthermore be integrated as part of an automated reduction and analysis procedure in laboratory instruments or at synchrotron beamlines.
5 Dec 2014
AG-2014.11-3559
physics.data-an
Hao Wu, Antonia S. J. S. Mey, Edina Rosta, Frank Noé
We propose a discrete transition-based reweighting analysis method (dTRAM) for analyzing configuration-space-discretized simulation trajectories produced at different thermodynamic states (temperatures, Hamiltonians, etc.) dTRAM provides maximum-likelihood estimates of stationary quantities (probabilities, free energies, expectation values) at any thermodynamic state. In contrast to the weighted histogram analysis method (WHAM), dTRAM does not require data to be sampled from global equilibrium, and can thus produce superior estimates for enhanced sampling data such as parallel/simulated tempering, replica exchange, umbrella sampling, or metadynamics. In addition, dTRAM provides optimal estimates of Markov state models (MSMs) from the discretized state-space trajectories at all thermodynamic states. Under suitable conditions, these MSMs can be used to calculate kinetic quantities (e.g. rates, timescales). In the limit of a single thermodynamic state, dTRAM estimates a maximum likelihood reversible MSM, while in the limit of uncorrelated sampling data, dTRAM is identical to WHAM. dTRAM is thus a generalization to both estimators.
30 Nov 2014
AG-2014.11-1588
physics.data-an
Aleksei Lokhov, Fyodor Tkachov
The method of quasi-optimal weights is applied to constructing (quasi-)optimal criteria for various anomalous contributions in experimental spectra. Anomalies in the spectra could indicate physics beyond the Standard Model (additional interactions and neutrino flavours, Lorenz violation etc.). In particular the cumulative tritium $β$-decay spectrum (for instance, in Troitsk-$ν$-mass, Mainz Neutrino Mass and KATRIN experiments) is analysed using the derived special criteria. Using the power functions we show that the derived quasi-optimal criteria are efficient statistical instruments for detecting the anomalous contributions in the spectra.
23 Nov 2014
AG-2014.11-2543
physics.data-an
Anton Poluektov
Kernel density estimation is a convenient way to estimate the probability density of a distribution given the sample of data points. However, it has certain drawbacks: proper description of the density using narrow kernels needs large data samples, whereas if the kernel width is large, boundaries and narrow structures tend to be smeared. Here, an approach to correct for such effects, is proposed that uses an approximate density to describe narrow structures and boundaries. The approach is shown to be well suited for the description of the efficiency shape over a multidimensional phase space in a typical particle physics analysis. An example is given for the five-dimensional phase space of the $Λ_b^0\to D^0pπ$ decay.
20 Nov 2014
AG-2014.11-1303
physics.data-an
Bernard Lacaze
In a communication scheme, there exist points at the transmitter and at the receiver where the wave is reduced to a finite set of functions of time which describe amplitudes and phases. For instance, the information is summarized in electrical cables which preceed or follow antennas. In many cases, a random propagation time is sufficient to explain changes induced by the medium. In this paper we study models based on stable probability laws which explain power spectra due to propagation of different kinds of waves in different media, for instance, acoustics in quiet or turbulent atmosphere, ultrasonics in liquids or tissues, or electromagnetic waves in free space or in cables. Physical examples show that a sub-class of probability laws appears in accordance with the causality property of linear filters.
19 Nov 2014
AG-2014.11-1316
physics.data-an
Jie Sun, Carlo Cafaro, Erik M. Bollt
Inferring the coupling structure of complex systems from time series data in general by means of statistical and information-theoretic techniques is a challenging problem in applied science. The reliability of statistical inferences requires the construction of suitable information-theoretic measures that take into account both direct and indirect influences, manifest in the form of information flows, between the components within the system. In this work, we present an application of the optimal causation entropy (oCSE) principle to identify the coupling structure of a synthetic biological system, the repressilator. Specifically, when the system reaches an equilibrium state, we use a stochastic perturbation approach to extract time series data that approximate a linear stochastic process. Then, we present and jointly apply the aggregative discovery and progressive removal algorithms based on the oCSE principle to infer the coupling structure of the system from the measured data. Finally, we show that the success rate of our coupling inferences not only improves with the amount of available data, but it also increases with a higher frequency of sampling and is especially immune to false positives.
19 Nov 2014
AG-2014.11-3711
physics.data-an
Alex Waagen, Raissa M. D'Souza
There is still much to discover about the mechanisms and nature of discontinuous percolation transitions. Much of the past work considers graph evolution algorithms known as Achlioptas processes in which a single edge is added to the graph from a set of $k$ randomly chosen candidate edges at each timestep until a giant component emerges. Several Achlioptas processes seem to yield a discontinuous percolation transition, but it was proven by Riordan and Warnke that the transition must be continuous in the thermodynamic limit. However, they also proved that if the number $k(n)$ of candidate edges increases with the number of nodes, then the percolation transition may be discontinuous. Here we attempt to find the simplest such process which yields a discontinuous transition in the thermodynamic limit. We introduce a process which considers only the degree of candidate edges and not component size. We calculate the critical point $t_{c}=(1-θ(\frac{1}{k}))n$ and rigorously show that the critical window is of size $O(\frac{n}{k(n)})$. If $k(n)$ grows very slowly, for example $k(n)=\log n$, the critical window is barely sublinear and hence the phase transition is discontinuous but appears continuous in finite systems. We also present arguments that Achlioptas processes with bounded size rules will always have continuous percolation transitions even with infinite choice.
17 Nov 2014
AG-2014.11-3658
physics.data-an
Araik Tamazian, Josef Ludescher, Armin Bunde
We study the distribution $P(x;α,L)$ of the relative trend $x$ in long-term correlated records of length $L$ that are characterized by a Hurst-exponent $α$ between 0.5 and 1.5 obtained by DFA2. The relative trend $x$ is the ratio between the strength of the trend $Δ$ in the record measured by linear regression, and the standard deviation $σ$ around the regression line. We consider $L$ between 400 and 2200, which is the typical length scale of monthly local and annual reconstructed global climate records. Extending previous work by Lennartz and Bunde \cite{Lennartz2011} we show explicitely that $P$ follows the student-t distribution $P\propto [1+(x/a)^2/l]^{-(l+1)/2}$, where the scaling parameter $a$ depends on both $L$ and $α$, while the effective length $l$ depends, for $α$ below 1.15, only on the record length $L$. From $P$ we can derive an analytical expression for the trend significance $S(x;α, L)=\int_{-x}^x P(x';α,L)dx'$ and the border lines of the $95\%$ percent significance interval. We show that the results are nearly independent of the distribution of the data in the record, holding for Gaussian data as well as for highly skewed non-Gaussian data. For an application, we use our methodology to estimate the significance of Central West Antarctic warming.
14 Nov 2014
AG-2014.11-880
physics.data-an
Anthony M. DeGennaro, Clarence W. Rowley, Luigi Martinelli
The formation and accretion of ice on the leading edge of a wing can be detrimental to airplane performance. Complicating this reality is the fact that even a small amount of uncertainty in the shape of the accreted ice may result in a large amount of uncertainty in aerodynamic performance metrics (e.g., stall angle of attack). The main focus of this work concerns using the techniques of Polynomial Chaos Expansions (PCE) to quantify icing uncertainty much more quickly than traditional methods (e.g., Monte Carlo). First, we present a brief survey of the literature concerning the physics of wing icing, with the intention of giving a certain amount of intuition for the physical process. Next, we give a brief overview of the background theory of PCE. Finally, we compare the results of Monte Carlo simulations to PCE-based uncertainty quantification for several different airfoil icing scenarios. The results are in good agreement and confirm that PCE methods are much more efficient for the canonical airfoil icing uncertainty quantification problem than Monte Carlo methods.
13 Nov 2014