ArXiv Daily GraphMining Perplexity Rules — Free Perplexity…
    Neura MarketNeura Market/Perplexity
    ChatGPTChatGPTClaudeClaudeGeminiGeminiCursorCursorGrokGrokPerplexityPerplexityDeepSeekDeepSeek
    CoPilotCoPilotStable DiffusionStable DiffusionMidjourneyMidjourney
    View All Directories
    OverviewRulesPromptsMCPsAgentsGamesBlogVideosGuidesCoursesCommunityTrending
    PerplexityRulesArXiv Daily GraphMining Perplexity Rules
    Back to Rules

    ArXiv Daily GraphMining Perplexity Rules

    alexfanjn July 19, 2026
    0 copies 0 downloads
    Rule Content
    [
      {
        "id": "arXiv:2212.01386",
        "title": "Convolution, aggregation and attention based deep neural networks for  accelerating simulations in mechanics",
        "abstract": "Deep learning surrogate models are being increasingly used in accelerating scientific simulations as a replacement for costly conventional numerical techniques. However, their use remains a significant challenge when dealing with real-world complex examples. In this work, we demonstrate three types of neural network architectures for efficient learning of highly non-linear deformations of solid bodies. The first two architectures are based on the recently proposed CNN U-NET and MAgNET (graph U-NET) frameworks which have shown promising performance for learning on mesh-based data. The third architecture is Perceiver IO, a very recent architecture that belongs to the family of attention-based neural networks--a class that has revolutionised diverse engineering fields and is still unexplored in computational mechanics. We study and compare the performance of all three networks on two benchmark examples, and show their capabilities to accurately predict the non-linear mechanical responses of soft bodies. ",
        "url": "https://arxiv.org/abs/2212.01386",
        "authors": [
          "Saurabh Deshpande",
          "Ra\u00fal I. Sosa",
          "St\u00e9phane P.A. Bordas",
          "Jakub Lengiewicz"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Computational Engineering, Finance, and Science (cs.CE)"
        ]
      },
      {
        "id": "arXiv:2212.01387",
        "title": "FullBrain: a Social E-learning Platform",
        "abstract": "We present FullBrain, a social e-learning platform where students share and track their knowledge. FullBrain users can post notes, ask questions and share learning resources in dedicated course and concept spaces. We detail two components of FullBrain: a SIR system equipped with query autocomplete and query autosuggestion, and a Leaderboard module to improve user experience. We analyzed the day-to-day users' usage of the SIR system, measuring a time-to-complete a request below 0.11s, matching or exceeding our UX targets. Moreover, we performed stress tests which lead the way for more detailed analysis. Through a preliminary user study and log data analysis, we observe that 97% of the users' activity is directed to the top 4 positions in the leaderboard. ",
        "url": "https://arxiv.org/abs/2212.01387",
        "authors": [
          "Mirko Biasini",
          "Vittorio Carmignani",
          "Nicola Ferro",
          "Panagiotis Filianos",
          "Maria Maistro",
          "Giorgio Maria di Nunzio"
        ],
        "subjectives": [
          "Human-Computer Interaction (cs.HC)"
        ]
      },
      {
        "id": "arXiv:2212.01424",
        "title": "PROB: Probabilistic Objectness for Open World Object Detection",
        "abstract": "Open World Object Detection (OWOD) is a new and challenging computer vision task that bridges the gap between classic object detection (OD) benchmarks and object detection in the real world. In addition to detecting and classifying seen/labeled objects, OWOD algorithms are expected to detect novel/unknown objects - which can be classified and incrementally learned. In standard OD, object proposals not overlapping with a labeled object are automatically classified as background. Therefore, simply applying OD methods to OWOD fails as unknown objects would be predicted as background. The challenge of detecting unknown objects stems from the lack of supervision in distinguishing unknown objects and background object proposals. Previous OWOD methods have attempted to overcome this issue by generating supervision using pseudo-labeling - however, unknown object detection has remained low. Probabilistic/generative models may provide a solution for this challenge. Herein, we introduce a novel probabilistic framework for objectness estimation, where we alternate between probability distribution estimation and objectness likelihood maximization of known objects in the embedded feature space - ultimately allowing us to estimate the objectness probability of different proposals. The resulting Probabilistic Objectness transformer-based open-world detector, PROB, integrates our framework into traditional object detection models, adapting them for the open-world setting. Comprehensive experiments on OWOD benchmarks show that PROB outperforms all existing OWOD methods in both unknown object detection ($\\sim 2\\times$ unknown recall) and known object detection ($\\sim 10\\%$ mAP). Our code will be made available upon publication at https://github.com/orrzohar/PROB. ",
        "url": "https://arxiv.org/abs/2212.01424",
        "authors": [
          "Orr Zohar",
          "Kuan-Chieh Wang",
          "Serena Yeung"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Artificial Intelligence (cs.AI)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.01426",
        "title": "Learning a Pedestrian Social Behavior Dictionary",
        "abstract": "Understanding pedestrian behavior patterns is a key component to building autonomous agents that can navigate among humans. We seek a learned dictionary of pedestrian behavior to obtain a semantic description of pedestrian trajectories. Supervised methods for dictionary learning are impractical since pedestrian behaviors may be unknown a priori and the process of manually generating behavior labels is prohibitively time consuming. We instead utilize a novel, unsupervised framework to create a taxonomy of pedestrian behavior observed in a specific space. First, we learn a trajectory latent space that enables unsupervised clustering to create an interpretable pedestrian behavior dictionary. We show the utility of this dictionary for building pedestrian behavior maps to visualize space usage patterns and for computing the distributions of behaviors. We demonstrate a simple but effective trajectory prediction by conditioning on these behavior labels. While many trajectory analysis methods rely on RNNs or transformers, we develop a lightweight, low-parameter approach and show results comparable to SOTA on the ETH and UCY datasets. ",
        "url": "https://arxiv.org/abs/2212.01426",
        "authors": [
          "Faith Johnson",
          "Kristin Dana"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.01437",
        "title": "Identifying Heterogeneous Treatment Effects in Multiple Outcomes using  Joint Confidence Intervals",
        "abstract": "Heterogeneous treatment effects (HTEs) are commonly identified during randomized controlled trials (RCTs). Identifying subgroups of patients with similar treatment effects is of high interest in clinical research to advance precision medicine. Often, multiple clinical outcomes are measured during an RCT, each having a potentially heterogeneous effect. Recently there has been high interest in identifying subgroups from HTEs, however, there has been less focus on developing tools in settings where there are multiple outcomes. In this work, we propose a framework for partitioning the covariate space to identify subgroups across multiple outcomes based on the joint CIs. We test our algorithm on synthetic and semi-synthetic data where there are two outcomes, and demonstrate that our algorithm is able to capture the HTE in both outcomes simultaneously. ",
        "url": "https://arxiv.org/abs/2212.01437",
        "authors": [
          "Peniel N. Argaw",
          "Elizabeth Healey",
          "Isaac S. Kohane"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.01440",
        "title": "SMARTQUERY: An Active Learning Framework for Graph Neural Networks  through Hybrid Uncertainty Reduction",
        "abstract": "Graph neural networks have achieved significant success in representation learning. However, the performance gains come at a cost; acquiring comprehensive labeled data for training can be prohibitively expensive. Active learning mitigates this issue by searching the unexplored data space and prioritizing the selection of data to maximize model's performance gain. In this paper, we propose a novel method SMARTQUERY, a framework to learn a graph neural network with very few labeled nodes using a hybrid uncertainty reduction function. This is achieved using two key steps: (a) design a multi-stage active graph learning framework by exploiting diverse explicit graph information and (b) introduce label propagation to efficiently exploit known labels to assess the implicit embedding information. Using a comprehensive set of experiments on three network datasets, we demonstrate the competitive performance of our method against state-of-the-arts on very few labeled data (up to 5 labeled nodes per class). ",
        "url": "https://arxiv.org/abs/2212.01440",
        "authors": [
          "Xiaoting Li",
          "Yuhang Wu",
          "Vineeth Rakesh",
          "Yusan Lin",
          "Hao Yang",
          "Fei Wang"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.01447",
        "title": "Compound Tokens: Channel Fusion for Vision-Language Representation  Learning",
        "abstract": "We present an effective method for fusing visual-and-language representations for several question answering tasks including visual question answering and visual entailment. In contrast to prior works that concatenate unimodal representations or use only cross-attention, we compose multimodal representations via channel fusion. By fusing on the channels, the model is able to more effectively align the tokens compared to standard methods. These multimodal representations, which we call compound tokens are generated with cross-attention transformer layers. First, vision tokens are used as queries to retrieve compatible text tokens through cross-attention. We then chain the vision tokens and the queried text tokens along the channel dimension. We call the resulting representations compound tokens. A second group of compound tokens are generated using an analogous process where the text tokens serve as queries to the cross-attention layer. We concatenate all the compound tokens for further processing with multimodal encoder. We demonstrate the effectiveness of compound tokens using an encoder-decoder vision-language model trained end-to-end in the open-vocabulary setting. Compound Tokens achieve highly competitive performance across a range of question answering tasks including GQA, VQA2.0, and SNLI-VE. ",
        "url": "https://arxiv.org/abs/2212.01447",
        "authors": [
          "Maxwell Mbabilla Aladago",
          "AJ Piergiovanni"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.01462",
        "title": "Topic Modeling on Clinical Social Work Notes for Exploring Social  Determinants of Health Factors",
        "abstract": "Most research studying social determinants of health (SDoH) has focused on physician notes or structured elements of the electronic medical record (EMR). We hypothesize that clinical notes from social workers, whose role is to ameliorate social and economic factors, might provide a richer source of data on SDoH. We sought to perform topic modeling to identify robust topics of discussion within a large cohort of social work notes. We retrieved a diverse, deidentified corpus of 0.95 million clinical social work notes from 181,644 patients at the University of California, San Francisco. We used word frequency analysis and Latent Dirichlet Allocation (LDA) topic modeling analysis to characterize this corpus and identify potential topics of discussion. Word frequency analysis identified both medical and non-medical terms associated with specific ICD10 chapters. The LDA topic modeling analysis extracted 11 topics related to social determinants of health risk factors including financial status, abuse history, social support, risk of death, and mental health. In addition, the topic modeling approach captured the variation between different types of social work notes and across patients with different types of diseases or conditions. We demonstrated that social work notes contain rich, unique, and otherwise unobtainable information on an individual's SDoH. ",
        "url": "https://arxiv.org/abs/2212.01462",
        "authors": [
          "Shenghuan Sun",
          "Travis Zack",
          "Madhumita Sushil",
          "Atul J. Butte"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)"
        ]
      },
      {
        "id": "arXiv:2212.01470",
        "title": "Prediction of Scene Plausibility",
        "abstract": "Understanding the 3D world from 2D images involves more than detection and segmentation of the objects within the scene. It also includes the interpretation of the structure and arrangement of the scene elements. Such understanding is often rooted in recognizing the physical world and its limitations, and in prior knowledge as to how similar typical scenes are arranged. In this research we pose a new challenge for neural network (or other) scene understanding algorithms - can they distinguish between plausible and implausible scenes? Plausibility can be defined both in terms of physical properties and in terms of functional and typical arrangements. Hence, we define plausibility as the probability of encountering a given scene in the real physical world. We build a dataset of synthetic images containing both plausible and implausible scenes, and test the success of various vision models in the task of recognizing and understanding plausibility. ",
        "url": "https://arxiv.org/abs/2212.01470",
        "authors": [
          "Or Nachmias",
          "Ohad Fried",
          "Ariel Shamir"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.01471",
        "title": "Conditions for Estimation of Sensitivities of Voltage Magnitudes to  Complex Power Injections",
        "abstract": "Voltage phase angle measurements are often unavailable from sensors in distribution networks and transmission network boundaries. Therefore, this paper addresses the conditions for estimating sensitivities of voltage magnitudes with respect to complex (active and reactive) electric power injections based on sensor measurements. These sensitivities represent submatrices of the inverse power flow Jacobian. We extend previous results to show that the sensitivities of a bus voltage magnitude with respect to active power injections are unique and different from those with respect to reactive power. The classical Newton-Raphson power flow model is used to derive a novel representation of bus voltage magnitudes as an underdetermined linear operator of the active and reactive power injections; parameterized by the bus power factors. Two conditions that ensure the existence of unique complex power injections given voltage magnitudes are established for this underdetermined linear system, thereby compressing the solution space. The first is a sufficient condition based on the bus power factors. The second is a necessary and sufficient condition based on the system eigenvalues. We use matrix completion theory to develop estimation methods for recovering sensitivity matrices with varying levels of sensor availability. Simulations verify the results and demonstrate engineering use of the proposed methods. ",
        "url": "https://arxiv.org/abs/2212.01471",
        "authors": [
          "Samuel Talkington",
          "Daniel Turizo",
          "Santiago Grijalva",
          "Jorge Fernandez",
          "Daniel K. Molzahn"
        ],
        "subjectives": [
          "Systems and Control (eess.SY)",
          "Optimization and Control (math.OC)"
        ]
      },
      {
        "id": "arXiv:2212.01479",
        "title": "Detecting Outdated Code Element References in Software Repository  Documentation",
        "abstract": "Outdated documentation is a pervasive problem in software development, preventing effective use of software, and misleading users and developers alike. We posit that one possible reason why documentation becomes out of sync so easily is that developers are unaware of when their source code modifications render the documentation obsolete. Ensuring that the documentation is always in sync with the source code takes considerable effort, especially for large codebases. To address this situation, we propose an approach that can automatically detect code element references that survive in the documentation after all source code instances have been deleted. In this work, we analysed over 3,000 GitHub projects and found that most projects contain at least one outdated code element reference at some point in their history. We submitted GitHub issues to real-world projects containing outdated references detected by our approach, some of which have already led to documentation fixes. As an initiative toward keeping documentation in software repositories up-to-date, we have made our implementation available for developers to scan their GitHub projects for outdated code element references. ",
        "url": "https://arxiv.org/abs/2212.01479",
        "authors": [
          "Wen Siang Tan",
          "Markus Wagner",
          "Christoph Treude"
        ],
        "subjectives": [
          "Software Engineering (cs.SE)"
        ]
      },
      {
        "id": "arXiv:2212.01515",
        "title": "Orders Are Unwanted: Dynamic Deep Graph Convolutional Network for  Personality Detection",
        "abstract": "Predicting personality traits based on online posts has emerged as an important task in many fields such as social network analysis. One of the challenges of this task is assembling information from various posts into an overall profile for each user. While many previous solutions simply concatenate the posts into a long document and then encode the document by sequential or hierarchical models, they introduce unwarranted orders for the posts, which may mislead the models. In this paper, we propose a dynamic deep graph convolutional network (D-DGCN) to overcome the above limitation. Specifically, we design a learn-to-connect approach that adopts a dynamic multi-hop structure instead of a deterministic structure, and combine it with a DGCN module to automatically learn the connections between posts. The modules of post encoder, learn-to-connect, and DGCN are jointly trained in an end-to-end manner. Experimental results on the Kaggle and Pandora datasets show the superior performance of D-DGCN to state-of-the-art baselines. Our code is available at https://github.com/djz233/D-DGCN. ",
        "url": "https://arxiv.org/abs/2212.01515",
        "authors": [
          "Tao Yang",
          "Jinghao Deng",
          "Xiaojun Quan",
          "Qifan Wang"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)"
        ]
      },
      {
        "id": "arXiv:2212.01528",
        "title": "IDMS: Instance Depth for Multi-scale Monocular 3D Object Detection",
        "abstract": "Due to the lack of depth information of images and poor detection accuracy in monocular 3D object detection, we proposed the instance depth for multi-scale monocular 3D object detection method. Firstly, to enhance the model's processing ability for different scale targets, a multi-scale perception module based on dilated convolution is designed, and the depth features containing multi-scale information are re-refined from both spatial and channel directions considering the inconsistency between feature maps of different scales. Firstly, we designed a multi-scale perception module based on dilated convolution to enhance the model's processing ability for different scale targets. The depth features containing multi-scale information are re-refined from spatial and channel directions considering the inconsistency between feature maps of different scales. Secondly, so as to make the model obtain better 3D perception, this paper proposed to use the instance depth information as an auxiliary learning task to enhance the spatial depth feature of the 3D target and use the sparse instance depth to supervise the auxiliary task. Finally, by verifying the proposed algorithm on the KITTI test set and evaluation set, the experimental results show that compared with the baseline method, the proposed method improves by 5.27\\% in AP40 in the car category, effectively improving the detection performance of the monocular 3D object detection algorithm. ",
        "url": "https://arxiv.org/abs/2212.01528",
        "authors": [
          "Chao Hu",
          "Liqiang Zhu",
          "Weibing Qiu",
          "Weijie Wu"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.01529",
        "title": "Laplacian Convolutional Representation for Traffic Time Series  Imputation",
        "abstract": "Spatiotemporal traffic data imputation is of great significance in intelligent transportation systems and data-driven decision-making processes. To make an accurate reconstruction on partially observed traffic data, we assert the importance of characterizing both global and local trends in traffic time series. In the literature, substantial prior works have demonstrated the effectiveness of utilizing low-rankness property of traffic data by matrix/tensor completion models. In this study, we first introduce a Laplacian kernel to temporal regularization for characterizing local trends in traffic time series, which can be formulated in the form of circular convolution. Then, we develop a low-rank Laplacian convolutional representation (LCR) model by putting the nuclear norm of a circulant matrix and the Laplacian temporal regularization together, which is proved to meet a unified framework that takes a fast Fourier transform solution in a relatively low time complexity. Through extensive experiments on some traffic datasets, we demonstrate the superiority of LCR for imputing traffic time series of various time series behaviors (e.g., data noises and strong/weak periodicity). The proposed LCR model is an efficient and effective solution to large-scale traffic data imputation over the existing baseline models. The adapted datasets and Python implementation are publicly available at https://github.com/xinychen/transdim. ",
        "url": "https://arxiv.org/abs/2212.01529",
        "authors": [
          "Xinyu Chen",
          "Zhanhong Cheng",
          "Nicolas Saunier",
          "Lijun Sun"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Artificial Intelligence (cs.AI)"
        ]
      },
      {
        "id": "arXiv:2212.01538",
        "title": "Multi-resolution Monocular Depth Map Fusion by Self-supervised  Gradient-based Composition",
        "abstract": "Monocular depth estimation is a challenging problem on which deep neural networks have demonstrated great potential. However, depth maps predicted by existing deep models usually lack fine-grained details due to the convolution operations and the down-samplings in networks. We find that increasing input resolution is helpful to preserve more local details while the estimation at low resolution is more accurate globally. Therefore, we propose a novel depth map fusion module to combine the advantages of estimations with multi-resolution inputs. Instead of merging the low- and high-resolution estimations equally, we adopt the core idea of Poisson fusion, trying to implant the gradient domain of high-resolution depth into the low-resolution depth. While classic Poisson fusion requires a fusion mask as supervision, we propose a self-supervised framework based on guided image filtering. We demonstrate that this gradient-based composition performs much better at noisy immunity, compared with the state-of-the-art depth map fusion method. Our lightweight depth fusion is one-shot and runs in real-time, making our method 80X faster than a state-of-the-art depth fusion method. Quantitative evaluations demonstrate that the proposed method can be integrated into many fully convolutional monocular depth estimation backbones with a significant performance boost, leading to state-of-the-art results of detail enhancement on depth maps. ",
        "url": "https://arxiv.org/abs/2212.01538",
        "authors": [
          "Yaqiao Dai",
          "Renjiao Yi",
          "Chenyang Zhu",
          "Hongjun He",
          "Kai Xu"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.01544",
        "title": "Probabilistic Verification of ReLU Neural Networks via Characteristic  Functions",
        "abstract": "Verifying the input-output relationships of a neural network so as to achieve some desired performance specification is a difficult, yet important, problem due to the growing ubiquity of neural nets in many engineering applications. We use ideas from probability theory in the frequency domain to provide probabilistic verification guarantees for ReLU neural networks. Specifically, we interpret a (deep) feedforward neural network as a discrete dynamical system over a finite horizon that shapes distributions of initial states, and use characteristic functions to propagate the distribution of the input data through the network. Using the inverse Fourier transform, we obtain the corresponding cumulative distribution function of the output set, which can be used to check if the network is performing as expected given any random point from the input set. The proposed approach does not require distributions to have well-defined moments or moment generating functions. We demonstrate our proposed approach on two examples, and compare its performance to related approaches. ",
        "url": "https://arxiv.org/abs/2212.01544",
        "authors": [
          "Joshua Pilipovsky",
          "Vignesh Sivaramakrishnan",
          "Meeko M. K. Oishi",
          "Panagiotis Tsiotras"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Systems and Control (eess.SY)",
          "Optimization and Control (math.OC)"
        ]
      },
      {
        "id": "arXiv:2212.01545",
        "title": "A Generalized Scalarization Method for Evolutionary Multi-objective  Optimization",
        "abstract": "The decomposition-based multi-objective evolutionary algorithm (MOEA/D) transforms a multi-objective optimization problem (MOP) into a set of single-objective subproblems for collaborative optimization. Mismatches between subproblems and solutions can lead to severe performance degradation of MOEA/D. Most existing mismatch coping strategies only work when the $L_{\\infty}$ scalarization is used. A mismatch coping strategy that can use any $L_{p}$ scalarization, even when facing MOPs with non-convex Pareto fronts, is of great significance for MOEA/D. This paper uses the global replacement (GR) as the backbone. We analyze how GR can no longer avoid mismatches when $L_{\\infty}$ is replaced by another $L_{p}$ with $p\\in [1,\\infty)$, and find that the $L_p$-based ($1\\leq p<\\infty$) subproblems having inconsistently large preference regions. When $p$ is set to a small value, some middle subproblems have very small preference regions so that their direction vectors cannot pass through their corresponding preference regions. Therefore, we propose a generalized $L_p$ (G$L_p$) scalarization to ensure that the subproblem's direction vector passes through its preference region. Our theoretical analysis shows that GR can always avoid mismatches when using the G$L_p$ scalarization for any $p\\geq 1$. The experimental studies on various MOPs conform to the theoretical analysis. ",
        "url": "https://arxiv.org/abs/2212.01545",
        "authors": [
          "Ruihao Zheng",
          "Zhenkun Wang"
        ],
        "subjectives": [
          "Neural and Evolutionary Computing (cs.NE)"
        ]
      },
      {
        "id": "arXiv:2212.01551",
        "title": "The Cause of Causal Emergence: Redistribution of Uncertainty",
        "abstract": "It is crucial to choose the appropriate scale in order to build an effective and informational representation of a complex system. Scientists carefully choose the scales for their experiments to extract the variables that describe the causalities in the system. They found that the coarse scale(macro) is sometimes more causal and informative than the numerous-parameter observations(micro). The phenomenon that the causality emerges by coarse-graining is called Causal Emergence(CE). Based on information theory, a number of recent works quantitatively showed that CE indeed happens while coarse-graining a micro model to the macro. However, the existing works have not discussed the question of why and when the CE happens. We quantitatively analyze the redistribution of uncertainties for coarse-graining and suggest that the redistribution of uncertainties is the cause of causal emergence. We further analyze the thresholds that determine if CE happens or not. From the regularity of the transition probability matrix(TPM) of discrete systems, the mathematical expressions of the model properties are derived. The values of thresholds for different operations are computed. The results provide the critical and specific conditions of CE as helpful suggestions for choosing the proper coarse-graining operation. The results also provided a new way to better understand the nature of causality and causal emergence. ",
        "url": "https://arxiv.org/abs/2212.01551",
        "authors": [
          "Liye Jia",
          "Cong Zhou",
          "Ka Lok Man",
          "Sheng-Uei Guan",
          "Jeremy Smith",
          "Yutao Yue"
        ],
        "subjectives": [
          "Information Theory (cs.IT)",
          "Artificial Intelligence (cs.AI)",
          "Computational Physics (physics.comp-ph)"
        ]
      },
      {
        "id": "arXiv:2212.01562",
        "title": "Understanding the Robustness of Multi-Exit Models under Common  Corruptions",
        "abstract": "Multi-Exit models (MEMs) use an early-exit strategy to improve the accuracy and efficiency of deep neural networks (DNNs) by allowing samples to exit the network before the last layer. However, the effectiveness of MEMs in the presence of distribution shifts remains largely unexplored. Our work examines how distribution shifts generated by common image corruptions affect the accuracy/efficiency of MEMs. We find that under common corruptions, early-exiting at the first correct exit reduces the inference cost and provides a significant boost in accuracy ( 10%) over exiting at the last layer. However, with realistic early-exit strategies, which do not assume knowledge about the correct exits, MEMs still reduce inference cost but provide a marginal improvement in accuracy (1%) compared to exiting at the last layer. Moreover, the presence of distribution shift widens the gap between an MEM's maximum classification accuracy and realistic early-exit strategies by 5% on average compared with the gap on in-distribution data. Our empirical analysis shows that the lack of calibration due to a distribution shift increases the susceptibility of such early-exit strategies to exit early and increases misclassification rates. Furthermore, the lack of calibration increases the inconsistency in the predictions of the model across exits, leading to both inefficient inference and more misclassifications compared with evaluation on in-distribution data. Finally, we propose two metrics, underthinking and overthinking, that quantify the different behavior of practical early-exit strategy under distribution shifts, and provide insights into improving the practical utility of MEMs. ",
        "url": "https://arxiv.org/abs/2212.01562",
        "authors": [
          "Akshay Mehra",
          "Skyler Seto",
          "Navdeep Jaitly",
          "Barry-John Theobald"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.01564",
        "title": "Node and Edge Centrality based Failures in Multi-layer Complex Networks",
        "abstract": "Multi-layer complex networks (MLCN) appears in various domains, such as, transportation, supply chains, etc. Failures in MLCN can lead to major disruptions in systems. Several research have focussed on different kinds of failures, such as, cascades, their reasons and ways to avoid them. This paper considers failures in a specific type of MLCN where the lower layer provides services to the higher layer without cross layer interaction, typical of a computer network. A three layer MLCN is constructed with the same set of nodes where each layer has different characteristics, the bottom most layer is Erdos-Renyi (ER) random graph with shortest path hop count among the nodes as gaussian, the middle layer is ER graph with higher number of edges from the previous, and the top most layer is scale free graph with even higher number of edges. Both edge and node failures are considered. Failures happen with decreasing order of centralities of edges and nodes in static batch mode and when the centralities change dynamically with progressive failures. Emergent pattern of three key parameters, namely, average shortest path length (ASPL), total shortest path count (TSPC) and total number of edges (TNE) for all the three layers after node or edge failures are studied. Extensive simulations show that all but one parameters show definite degrading patterns. Surprising, ASPL for the middle layer starts showing a chaotic behavior beyond a certain point for all types of failures. ",
        "url": "https://arxiv.org/abs/2212.01564",
        "authors": [
          "Dibakar Das",
          "Jyotsna Bapat",
          "Debabrata Das"
        ],
        "subjectives": [
          "Social and Information Networks (cs.SI)",
          "Physics and Society (physics.soc-ph)"
        ]
      },
      {
        "id": "arXiv:2212.01565",
        "title": "Leveraging Angular Information Between Feature and Classifier for  Long-tailed Learning: A Prediction Reformulation Approach",
        "abstract": "Deep neural networks still struggle on long-tailed image datasets, and one of the reasons is that the imbalance of training data across categories leads to the imbalance of trained model parameters. Motivated by the empirical findings that trained classifiers yield larger weight norms in head classes, we propose to reformulate the recognition probabilities through included angles without re-balancing the classifier weights. Specifically, we calculate the angles between the data feature and the class-wise classifier weights to obtain angle-based prediction results. Inspired by the performance improvement of the predictive form reformulation and the outstanding performance of the widely used two-stage learning framework, we explore the different properties of this angular prediction and propose novel modules to improve the performance of different components in the framework. Our method is able to obtain the best performance among peer methods without pretraining on CIFAR10/100-LT and ImageNet-LT. Source code will be made publicly available. ",
        "url": "https://arxiv.org/abs/2212.01565",
        "authors": [
          "Haoxuan Wang",
          "Junchi Yan"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.01568",
        "title": "Generalizing Multiple Object Tracking to Unseen Domains by Introducing  Natural Language Representation",
        "abstract": "Although existing multi-object tracking (MOT) algorithms have obtained competitive performance on various benchmarks, almost all of them train and validate models on the same domain. The domain generalization problem of MOT is hardly studied. To bridge this gap, we first draw the observation that the high-level information contained in natural language is domain invariant to different tracking domains. Based on this observation, we propose to introduce natural language representation into visual MOT models for boosting the domain generalization ability. However, it is infeasible to label every tracking target with a textual description. To tackle this problem, we design two modules, namely visual context prompting (VCP) and visual-language mixing (VLM). Specifically, VCP generates visual prompts based on the input frames. VLM joints the information in the generated visual prompts and the textual prompts from a pre-defined Trackbook to obtain instance-level pseudo textual description, which is domain invariant to different tracking scenes. Through training models on MOT17 and validating them on MOT20, we observe that the pseudo textual descriptions generated by our proposed modules improve the generalization performance of query-based trackers by large margins. ",
        "url": "https://arxiv.org/abs/2212.01568",
        "authors": [
          "En Yu",
          "Songtao Liu",
          "Zhuoling Li",
          "Jinrong Yang",
          "Zeming li",
          "Shoudong Han",
          "Wenbing Tao"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.01581",
        "title": "Modeling Label Correlations for Ultra-Fine Entity Typing with Neural  Pairwise Conditional Random Field",
        "abstract": "Ultra-fine entity typing (UFET) aims to predict a wide range of type phrases that correctly describe the categories of a given entity mention in a sentence. Most recent works infer each entity type independently, ignoring the correlations between types, e.g., when an entity is inferred as a president, it should also be a politician and a leader. To this end, we use an undirected graphical model called pairwise conditional random field (PCRF) to formulate the UFET problem, in which the type variables are not only unarily influenced by the input but also pairwisely relate to all the other type variables. We use various modern backbones for entity typing to compute unary potentials, and derive pairwise potentials from type phrase representations that both capture prior semantic information and facilitate accelerated inference. We use mean-field variational inference for efficient type inference on very large type sets and unfold it as a neural network module to enable end-to-end training. Experiments on UFET show that the Neural-PCRF consistently outperforms its backbones with little cost and results in a competitive performance against cross-encoder based SOTA while being thousands of times faster. We also find Neural- PCRF effective on a widely used fine-grained entity typing dataset with a smaller type set. We pack Neural-PCRF as a network module that can be plugged onto multi-label type classifiers with ease and release it in https://github.com/modelscope/adaseq/tree/master/examples/NPCRF. ",
        "url": "https://arxiv.org/abs/2212.01581",
        "authors": [
          "Chengyue Jiang",
          "Yong Jiang",
          "Weiqi Wu",
          "Pengjun Xie",
          "Kewei Tu"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)"
        ]
      },
      {
        "id": "arXiv:2212.01602",
        "title": "StegaNeRF: Embedding Invisible Information within Neural Radiance Fields",
        "abstract": "Recent advances in neural rendering imply a future of widespread visual data distributions through sharing NeRF model weights. However, while common visual data (images and videos) have standard approaches to embed ownership or copyright information explicitly or subtly, the problem remains unexplored for the emerging NeRF format. We present StegaNeRF, a method for steganographic information embedding in NeRF renderings. We design an optimization framework allowing accurate hidden information extractions from images rendered by NeRF, while preserving its original visual quality. We perform experimental evaluations of our method under several potential deployment scenarios, and we further discuss the insights discovered through our analysis. StegaNeRF signifies an initial exploration into the novel problem of instilling customizable, imperceptible, and recoverable information to NeRF renderings, with minimal impact to rendered images. Project page: https://xggnet.github.io/StegaNeRF/. ",
        "url": "https://arxiv.org/abs/2212.01602",
        "authors": [
          "Chenxin Li",
          "Brandon Y. Feng",
          "Zhiwen Fan",
          "Panwang Pan",
          "Zhangyang Wang"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.01606",
        "title": "An ADMM-Incorporated Latent Factorization of Tensors Method for QoS  Prediction",
        "abstract": "As the Internet developed rapidly, it is important to choose suitable web services from a wide range of candidates. Quality of service (QoS) describes the performance of a web service dynamically with respect to the service requested by the service consumer. Moreover, the latent factorization of tenors (LFT) is very effective for discovering temporal patterns in high dimensional and sparse (HiDS) tensors. However, current LFT models suffer from a low convergence rate and rarely account for the effects of outliers. To address the above problems, this paper proposes an Alternating direction method of multipliers (ADMM)-based Outlier-Resilient Nonnegative Latent-factorization of Tensors model. We maintain the non-negativity of the model by constructing an augmented Lagrangian function with the ADMM optimization framework. In addition, the Cauchy function is taken as the metric function to reduce the impact on the model training. The empirical work on two dynamic QoS datasets shows that the proposed method has faster convergence and better performance on prediction accuracy. ",
        "url": "https://arxiv.org/abs/2212.01606",
        "authors": [
          "Jiajia Mi",
          "Hao Wu"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.01610",
        "title": "Exploring Stochastic Autoregressive Image Modeling for Visual  Representation",
        "abstract": "Autoregressive language modeling (ALM) have been successfully used in self-supervised pre-training in Natural language processing (NLP). However, this paradigm has not achieved comparable results with other self-supervised approach in computer vision (e.g., contrastive learning, mask image modeling). In this paper, we try to find the reason why autoregressive modeling does not work well on vision tasks. To tackle this problem, we fully analyze the limitation of visual autoregressive methods and proposed a novel stochastic autoregressive image modeling (named SAIM) by the two simple designs. First, we employ stochastic permutation strategy to generate effective and robust image context which is critical for vision tasks. Second, we create a parallel encoder-decoder training process in which the encoder serves a similar role to the standard vision transformer focus on learning the whole contextual information, and meanwhile the decoder predicts the content of the current position, so that the encoder and decoder can reinforce each other. By introducing stochastic prediction and the parallel encoder-decoder, SAIM significantly improve the performance of autoregressive image modeling. Our method achieves the best accuracy (83.9%) on the vanilla ViT-Base model among methods using only ImageNet-1K data. Transfer performance in downstream tasks also show that our model achieves competitive performance. ",
        "url": "https://arxiv.org/abs/2212.01610",
        "authors": [
          "Yu Qi",
          "Fan Yang",
          "Yousong Zhu",
          "Yufei Liu",
          "Liwei Wu",
          "Rui Zhao",
          "Wei Li"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Artificial Intelligence (cs.AI)"
        ]
      },
      {
        "id": "arXiv:2212.01611",
        "title": "CoP: Factual Inconsistency Detection by Controlling the Preference",
        "abstract": "Abstractive summarization is the process of generating a summary given a document as input. Although significant progress has been made, the factual inconsistency between the document and the generated summary still limits its practical applications. Previous work found that the probabilities assigned by the generation model reflect its preferences for the generated summary, including the preference for factual consistency, and the preference for the language or knowledge prior as well. To separate the preference for factual consistency, we propose an unsupervised framework named CoP by controlling the preference of the generation model with the help of prompt. More specifically, the framework performs an extra inference step in which a text prompt is introduced as an additional input. In this way, another preference is described by the generation probability of this extra inference process. The difference between the above two preferences, i.e. the difference between the probabilities, could be used as measurements for detecting factual inconsistencies. Interestingly, we found that with the properly designed prompt, our framework could evaluate specific preferences and serve as measurements for fine-grained categories of inconsistency, such as entity-related inconsistency, coreference-related inconsistency, etc. Moreover, our framework could also be extended to the supervised setting to learn better prompt from the labeled data as well. Experiments show that our framework achieves new SOTA results on three factual inconsistency detection tasks. ",
        "url": "https://arxiv.org/abs/2212.01611",
        "authors": [
          "Shuaijie She",
          "Xiang Geng",
          "Shujian Huang",
          "Jiajun Chen"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)"
        ]
      },
      {
        "id": "arXiv:2212.01614",
        "title": "On the Performance of Non-Terrestrial Networks to Support the Internet  of Things",
        "abstract": "The advent of the Internet of Things (IoT) era, where billions of devices and sensors are becoming more and more connected and ubiquitous, is putting a strain on traditional terrestrial networks, that may no longer be able to fulfill service requirements efficiently. This issue is further complicated in rural and remote areas with scarce and low-quality cellular coverage. To fill this gap, the research community is focusing on non-terrestrial networks (NTNs), where Unmanned Aerial Vehicles (UAVs), High Altitude Platforms (HAPs) and satellites can serve as aerial/space gateways to aggregate, process, and relay the IoT traffic. In this paper we demonstrate this paradigm, and evaluate how common Low-Power Wide Area Network (LPWAN) technologies, designed and developed to operate for IoT systems, work in NTNs. We then formalize an optimization problem to decide whether and how IoT traffic can be offloaded to LEO satellites to reduce the burden on terrestrial gateways. ",
        "url": "https://arxiv.org/abs/2212.01614",
        "authors": [
          "Dengke Wang",
          "Alessandro Traspadini",
          "Marco Giordani",
          "Mohamed-Slim Alouini",
          "Michele Zorzi"
        ],
        "subjectives": [
          "Networking and Internet Architecture (cs.NI)"
        ]
      },
      {
        "id": "arXiv:2212.01618",
        "title": "An Overview of Trust Standards for Communication Networks and Future  Digital World",
        "abstract": "With the development of Information and Communication Technologies, trust has been applied more and more in various scenarios. At the same time, different organizations have published a series of trust frameworks to support the implementation of trust. There are also academic paper discussing about these trust standards, however, most of them only focus on a specific application. Unlike existing works, this paper provides an overview of all current available trust standards related to communication networks and future digital world from several main organizations. To be specific, this paper summarizes and organizes all these trust standards into three layers: trust foundation, trust elements, and trust applications. We then analysis these trust standards and discuss their contribution in a systematic way. We discuss the motivations behind each current in forced standards, analyzes their frameworks and solutions, and presents their role and impact on communication works and future digital world. Finally, we give our suggestions on the trust work that needs to be standardized in future. ",
        "url": "https://arxiv.org/abs/2212.01618",
        "authors": [
          "Huilin Wang",
          "Xin Kang",
          "Tieyan Li",
          "Zhongding Lei",
          "Cheng-Kang Chu",
          "Haiguang Wang"
        ],
        "subjectives": [
          "Information Theory (cs.IT)",
          "Cryptography and Security (cs.CR)"
        ]
      },
      {
        "id": "arXiv:2212.01627",
        "title": "Castell: Scalable Joint Probability Estimation of Multi-dimensional Data  Randomized with Local Differential Privacy",
        "abstract": "Performing randomized response (RR) over multi-dimensional data is subject to the curse of dimensionality. As the number of attributes increases, the exponential growth in the number of attribute-value combinations greatly impacts the computational cost and the accuracy of the RR estimates. In this paper, we propose a new multi-dimensional RR scheme that randomizes all attributes independently, and then aggregates these randomization matrices into a single aggregated matrix. The multi-dimensional joint probability distributions are then estimated. The inverse matrix of the aggregated randomization matrix can be computed efficiently at a lightweight computation cost (i.e., linear with respect to dimensionality) and with manageable storage requirements. To overcome the limitation of accuracy, we propose two extensions to the baseline protocol, called {\\em hybrid} and {\\em truncated} schemes. Finally, we have conducted experiments using synthetic and major open-source datasets for various numbers of attributes, domain sizes, and numbers of respondents. The results using UCI Adult dataset give average distances between the estimated and the real (2 through 6-way) joint probability are $0.0099$ for {\\em truncated} and $0.0155$ for {\\em hybrid} schemes, whereas they are $0.03$ and $0.04$ for LoPub, which is the state-of-the-art multi-dimensional LDP scheme. ",
        "url": "https://arxiv.org/abs/2212.01627",
        "authors": [
          "Hiroaki Kikuchi"
        ],
        "subjectives": [
          "Cryptography and Security (cs.CR)"
        ]
      },
      {
        "id": "arXiv:2212.01629",
        "title": "Generating Synthetic Data in a Secure Federated General Adversarial  Networks for a Consortium of Health Registries",
        "abstract": "In this work, we review the architecture design of existing federated General Adversarial Networks (GAN) solutions and highlight the security and trust-related weaknesses in the existing designs. We then describe how these weaknesses make existing designs unsuitable for the requirements needed for a consortium of health registries working towards generating synthetic datasets for research purposes. Moreover, we propose how these weaknesses can be addressed with our novel architecture solution. Our novel architecture solution combines several building blocks to generate synthetic data in a decentralised setting. Consortium blockchains, secure multi-party computations, and homomorphic encryption are the core building blocks of our proposed architecture solution to address the weaknesses in the existing design of federated GANs. Finally, we discuss our proposed solution's advantages and future research directions. ",
        "url": "https://arxiv.org/abs/2212.01629",
        "authors": [
          "Narasimha Raghavan Veeraragavan",
          "Jan Franz Nyg\u00e5rd"
        ],
        "subjectives": [
          "Distributed, Parallel, and Cluster Computing (cs.DC)"
        ]
      },
      {
        "id": "arXiv:2212.01641",
        "title": "Intermediate Entity-based Sparse Interpretable Representation Learning",
        "abstract": "Interpretable entity representations (IERs) are sparse embeddings that are \"human-readable\" in that dimensions correspond to fine-grained entity types and values are predicted probabilities that a given entity is of the corresponding type. These methods perform well in zero-shot and low supervision settings. Compared to standard dense neural embeddings, such interpretable representations may permit analysis and debugging. However, while fine-tuning sparse, interpretable representations improves accuracy on downstream tasks, it destroys the semantics of the dimensions which were enforced in pre-training. Can we maintain the interpretable semantics afforded by IERs while improving predictive performance on downstream tasks? Toward this end, we propose Intermediate enTity-based Sparse Interpretable Representation Learning (ItsIRL). ItsIRL realizes improved performance over prior IERs on biomedical tasks, while maintaining \"interpretability\" generally and their ability to support model debugging specifically. The latter is enabled in part by the ability to perform \"counterfactual\" fine-grained entity type manipulation, which we explore in this work. Finally, we propose a method to construct entity type based class prototypes for revealing global semantic properties of classes learned by our model. ",
        "url": "https://arxiv.org/abs/2212.01641",
        "authors": [
          "Diego Garcia-Olano",
          "Yasumasa Onoe",
          "Joydeep Ghosh",
          "Byron C. Wallace"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.01662",
        "title": "Modeling Mobile Visualization for Medical Reports of Complex Chronic  Diseases",
        "abstract": "Visualizing medical histories of patients with complex chronic diseases (e.g., discordant chronic comorbidities (DCCs)) is a challenge for patients, their healthcare providers, and their support network. DCCs are health conditions in which patients have multiple, often unrelated, chronic illnesses that may need to be addressed concurrently but may also be associated with conflicting treatment instructions. Future work targeting to reduce treatment conflicts and improve patient quality of life and care should carefully examine and visualize DCCs medical reports, symptoms, and treatment recommendations. In this study, we explore various visualization models and paradigms. We analyze how these models and paradigms are applied to visualize multifaceted medical data. We then propose a model for transforming the unstructured data into temporal slices and depict them in a single graphic model. We report how we carefully moved multifaceted DCC records into; structured data tables, visualization graphs, and various hardware devices. ",
        "url": "https://arxiv.org/abs/2212.01662",
        "authors": [
          "Sankarshan Dasgupta",
          "Tom Ongwere"
        ],
        "subjectives": [
          "Human-Computer Interaction (cs.HC)",
          "Multimedia (cs.MM)"
        ]
      },
      {
        "id": "arXiv:2212.01667",
        "title": "T-STAR: Truthful Style Transfer using AMR Graph as Intermediate  Representation",
        "abstract": "Unavailability of parallel corpora for training text style transfer (TST) models is a very challenging yet common scenario. Also, TST models implicitly need to preserve the content while transforming a source sentence into the target style. To tackle these problems, an intermediate representation is often constructed that is devoid of style while still preserving the meaning of the source sentence. In this work, we study the usefulness of Abstract Meaning Representation (AMR) graph as the intermediate style agnostic representation. We posit that semantic notations like AMR are a natural choice for an intermediate representation. Hence, we propose T-STAR: a model comprising of two components, text-to-AMR encoder and a AMR-to-text decoder. We propose several modeling improvements to enhance the style agnosticity of the generated AMR. To the best of our knowledge, T-STAR is the first work that uses AMR as an intermediate representation for TST. With thorough experimental evaluation we show T-STAR significantly outperforms state of the art techniques by achieving on an average 15.2% higher content preservation with negligible loss (3% approx.) in style accuracy. Through detailed human evaluation with 90,000 ratings, we also show that T-STAR has up to 50% lesser hallucinations compared to state of the art TST models. ",
        "url": "https://arxiv.org/abs/2212.01667",
        "authors": [
          "Anubhav Jangra",
          "Preksha Nema",
          "Aravindan Raghuveer"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)"
        ]
      },
      {
        "id": "arXiv:2212.01672",
        "title": "MaRF: Representing Mars as Neural Radiance Fields",
        "abstract": "The aim of this work is to introduce MaRF, a novel framework able to synthesize the Martian environment using several collections of images from rover cameras. The idea is to generate a 3D scene of Mars' surface to address key challenges in planetary surface exploration such as: planetary geology, simulated navigation and shape analysis. Although there exist different methods to enable a 3D reconstruction of Mars' surface, they rely on classical computer graphics techniques that incur high amounts of computational resources during the reconstruction process, and have limitations with generalizing reconstructions to unseen scenes and adapting to new images coming from rover cameras. The proposed framework solves the aforementioned limitations by exploiting Neural Radiance Fields (NeRFs), a method that synthesize complex scenes by optimizing a continuous volumetric scene function using a sparse set of images. To speed up the learning process, we replaced the sparse set of rover images with their neural graphics primitives (NGPs), a set of vectors of fixed length that are learned to preserve the information of the original images in a significantly smaller size. In the experimental section, we demonstrate the environments created from actual Mars datasets captured by Curiosity rover, Perseverance rover and Ingenuity helicopter, all of which are available on the Planetary Data System (PDS). ",
        "url": "https://arxiv.org/abs/2212.01672",
        "authors": [
          "Lorenzo Giusti",
          "Josue Garcia",
          "Steven Cozine",
          "Darrick Suen",
          "Christina Nguyen",
          "Ryan Alimo"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Graphics (cs.GR)"
        ]
      },
      {
        "id": "arXiv:2212.01682",
        "title": "Interpretable Node Representation with Attribute Decoding",
        "abstract": "Variational Graph Autoencoders (VGAEs) are powerful models for unsupervised learning of node representations from graph data. In this work, we systematically analyze modeling node attributes in VGAEs and show that attribute decoding is important for node representation learning. We further propose a new learning model, interpretable NOde Representation with Attribute Decoding (NORAD). The model encodes node representations in an interpretable approach: node representations capture community structures in the graph and the relationship between communities and node attributes. We further propose a rectifying procedure to refine node representations of isolated notes, improving the quality of these nodes' representations. Our empirical results demonstrate the advantage of the proposed model when learning graph data in an interpretable approach. ",
        "url": "https://arxiv.org/abs/2212.01682",
        "authors": [
          "Xiaohui Chen",
          "Xi Chen",
          "Liping Liu"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Social and Information Networks (cs.SI)"
        ]
      },
      {
        "id": "arXiv:2212.01683",
        "title": "Recognition and Prediction of Surgical Gestures and Trajectories Using  Transformer Models in Robot-Assisted Surgery",
        "abstract": "Surgical activity recognition and prediction can help provide important context in many Robot-Assisted Surgery (RAS) applications, for example, surgical progress monitoring and estimation, surgical skill evaluation, and shared control strategies during teleoperation. Transformer models were first developed for Natural Language Processing (NLP) to model word sequences and soon the method gained popularity for general sequence modeling tasks. In this paper, we propose the novel use of a Transformer model for three tasks: gesture recognition, gesture prediction, and trajectory prediction during RAS. We modify the original Transformer architecture to be able to generate the current gesture sequence, future gesture sequence, and future trajectory sequence estimations using only the current kinematic data of the surgical robot end-effectors. We evaluate our proposed models on the JHU-ISI Gesture and Skill Assessment Working Set (JIGSAWS) and use Leave-One-User-Out (LOUO) cross-validation to ensure the generalizability of our results. Our models achieve up to 89.3\\% gesture recognition accuracy, 84.6\\% gesture prediction accuracy (1 second ahead) and 2.71mm trajectory prediction error (1 second ahead). Our models are comparable to and able to outperform state-of-the-art methods while using only the kinematic data channel. This approach can enable near-real time surgical activity recognition and prediction. ",
        "url": "https://arxiv.org/abs/2212.01683",
        "authors": [
          "Chang Shi",
          "Yi Zheng",
          "Ann Majewicz Fey"
        ],
        "subjectives": [
          "Robotics (cs.RO)",
          "Artificial Intelligence (cs.AI)"
        ]
      },
      {
        "id": "arXiv:2212.01688",
        "title": "LDL: A Defense for Label-Based Membership Inference Attacks",
        "abstract": "The data used to train deep neural network (DNN) models in applications such as healthcare and finance typically contain sensitive information. A DNN model may suffer from overfitting. Overfitted models have been shown to be susceptible to query-based attacks such as membership inference attacks (MIAs). MIAs aim to determine whether a sample belongs to the dataset used to train a classifier (members) or not (nonmembers). Recently, a new class of label based MIAs (LAB MIAs) was proposed, where an adversary was only required to have knowledge of predicted labels of samples. Developing a defense against an adversary carrying out a LAB MIA on DNN models that cannot be retrained remains an open problem. We present LDL, a light weight defense against LAB MIAs. LDL works by constructing a high-dimensional sphere around queried samples such that the model decision is unchanged for (noisy) variants of the sample within the sphere. This sphere of label-invariance creates ambiguity and prevents a querying adversary from correctly determining whether a sample is a member or a nonmember. We analytically characterize the success rate of an adversary carrying out a LAB MIA when LDL is deployed, and show that the formulation is consistent with experimental observations. We evaluate LDL on seven datasets -- CIFAR-10, CIFAR-100, GTSRB, Face, Purchase, Location, and Texas -- with varying sizes of training data. All of these datasets have been used by SOTA LAB MIAs. Our experiments demonstrate that LDL reduces the success rate of an adversary carrying out a LAB MIA in each case. We empirically compare LDL with defenses against LAB MIAs that require retraining of DNN models, and show that LDL performs favorably despite not needing to retrain the DNNs. ",
        "url": "https://arxiv.org/abs/2212.01688",
        "authors": [
          "Arezoo Rajabi",
          "Dinuka Sahabandu",
          "Luyao Niu",
          "Bhaskar Ramasubramanian",
          "Radha Poovendran"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Cryptography and Security (cs.CR)"
        ]
      },
      {
        "id": "arXiv:2212.01698",
        "title": "Precise Energy Consumption Measurements of Heterogeneous Artificial  Intelligence Workloads",
        "abstract": "With the rise of AI in recent years and the increase in complexity of the models, the growing demand in computational resources is starting to pose a significant challenge. The need for higher compute power is being met with increasingly more potent accelerators and the use of large compute clusters. However, the gain in prediction accuracy from large models trained on distributed and accelerated systems comes at the price of a substantial increase in energy demand, and researchers have started questioning the environmental friendliness of such AI methods at scale. Consequently, energy efficiency plays an important role for AI model developers and infrastructure operators alike. The energy consumption of AI workloads depends on the model implementation and the utilized hardware. Therefore, accurate measurements of the power draw of AI workflows on different types of compute nodes is key to algorithmic improvements and the design of future compute clusters and hardware. To this end, we present measurements of the energy consumption of two typical applications of deep learning models on different types of compute nodes. Our results indicate that 1. deriving energy consumption directly from runtime is not accurate, but the consumption of the compute node needs to be considered regarding its composition; 2. neglecting accelerator hardware on mixed nodes results in overproportional inefficiency regarding energy consumption; 3. energy consumption of model training and inference should be considered separately - while training on GPUs outperforms all other node types regarding both runtime and energy consumption, inference on CPU nodes can be comparably efficient. One advantage of our approach is that the information on energy consumption is available to all users of the supercomputer, enabling an easy transfer to other workloads alongside a raise in user-awareness of energy consumption. ",
        "url": "https://arxiv.org/abs/2212.01698",
        "authors": [
          "Ren\u00e9 Caspart",
          "Sebastian Ziegler",
          "Arvid Weyrauch",
          "Holger Obermaier",
          "Simon Raffeiner",
          "Leon Pascal Schuhmacher",
          "Jan Scholtyssek",
          "Darya Trofimova",
          "Marco Nolden",
          "Ines Reinartz",
          "Fabian Isensee",
          "Markus G\u00f6tz",
          "Charlotte Debus"
        ],
        "subjectives": [
          "Distributed, Parallel, and Cluster Computing (cs.DC)",
          "Artificial Intelligence (cs.AI)"
        ]
      },
      {
        "id": "arXiv:2212.01700",
        "title": "Towards Robust NLG Bias Evaluation with Syntactically-diverse Prompts",
        "abstract": "We present a robust methodology for evaluating biases in natural language generation(NLG) systems. Previous works use fixed hand-crafted prefix templates with mentions of various demographic groups to prompt models to generate continuations for bias analysis. These fixed prefix templates could themselves be specific in terms of styles or linguistic structures, which may lead to unreliable fairness conclusions that are not representative of the general trends from tone varying prompts. To study this problem, we paraphrase the prompts with different syntactic structures and use these to evaluate demographic bias in NLG systems. Our results suggest similar overall bias trends but some syntactic structures lead to contradictory conclusions compared to past works. We show that our methodology is more robust and that some syntactic structures prompt more toxic content while others could prompt less biased generation. This suggests the importance of not relying on a fixed syntactic structure and using tone-invariant prompts. Introducing syntactically-diverse prompts can achieve more robust NLG (bias) evaluation. ",
        "url": "https://arxiv.org/abs/2212.01700",
        "authors": [
          "Arshiya Aggarwal",
          "Jiao Sun",
          "Nanyun Peng"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)"
        ]
      },
      {
        "id": "arXiv:2212.01701",
        "title": "Social Stratification in Networks: Insights from Co-Authorship Networks",
        "abstract": "It has been observed that real-world social networks often exhibit stratification along economic or other lines, with consequences for class mobility and access to opportunities. With the rise in human interaction data and extensive use of online social networks, the structure of social networks (representing connections between individuals) can be used for measuring stratification. However, although stratification has been studied extensively in the social sciences, there is no single, generally applicable metric for measuring the level of stratification in a network. In this work, we first propose the novel Stratification Assortativity (StA) metric, which measures the extent to which a network is stratified into different tiers. Then, we use the \\texttt{StA} metric to perform an in-depth analysis of the stratification of five co-authorship networks. We examine the evolution of these networks over 50 years and show that these fields demonstrate an increasing level of stratification over time, and, correspondingly, the trajectory of a researcher's career is increasingly correlated with her entry point into the network. ",
        "url": "https://arxiv.org/abs/2212.01701",
        "authors": [
          "Zeinab S. Jalali",
          "Josh Introne",
          "Sucheta Soundarajan"
        ],
        "subjectives": [
          "Social and Information Networks (cs.SI)"
        ]
      },
      {
        "id": "arXiv:2212.01711",
        "title": "Linguistic Constructs as the Representation of the Domain Model in an  Intelligent Language Tutoring System",
        "abstract": "This paper presents the development of an AI-based language learning platform Revita. It is a freely available intelligent online tutor, developed to support learners of multiple languages, from low-intermediate to advanced levels. It has been in pilot use by hundreds of students at several universities, whose feedback and needs are shaping the development. One of the main emerging features of Revita is the introduction of a system of linguistic constructs as the representation of domain knowledge. The system of constructs is developed in close collaboration with experts in language teaching. Constructs define the types of exercises, the content of the feedback, and enable the detailed modeling and evaluation of learning progress. ",
        "url": "https://arxiv.org/abs/2212.01711",
        "authors": [
          "Anisia Katinskaia",
          "Jue Hou",
          "Anh-Duc Vu",
          "Roman Yangarber"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)"
        ]
      },
      {
        "id": "arXiv:2212.01717",
        "title": "Variational Bayes for Joint Channel Estimation and Data Detection in  Few-Bit Massive MIMO Systems",
        "abstract": "Massive multiple-input multiple-output (MIMO) communications using low-resolution analog-to-digital converters (ADCs) is a promising technology for providing high spectral and energy efficiency with affordable hardware cost and power consumption. However, the use of low-resolution ADCs requires special signal processing methods for channel estimation and data detection since the resulting system is severely non-linear. This paper proposes joint channel estimation and data detection methods for massive MIMO systems with low-resolution ADCs based on the variational Bayes (VB) inference framework. We first derive matched-filter quantized VB (MF-QVB) and linear minimum mean-squared error quantized VB (LMMSE-QVB) detection methods assuming the channel state information (CSI) is available. Then we extend these methods to the joint channel estimation and data detection (JED) problem and propose two methods we refer to as MF-QVB-JED and LMMSE-QVB-JED. Unlike conventional VB-based detection methods that assume knowledge of the second-order statistics of the additive noise, we propose to float the noise variance/covariance matrix as an unknown random variable that is used to account for both the noise and the residual inter-user interference. We also present practical aspects of the QVB framework to improve its implementation stability. Finally, we show via numerical results that the proposed VB-based methods provide robust performance and also significantly outperform existing methods. ",
        "url": "https://arxiv.org/abs/2212.01717",
        "authors": [
          "Ly V. Nguyen",
          "A. Lee Swindlehurst",
          "Duy H. N. Nguyen"
        ],
        "subjectives": [
          "Information Theory (cs.IT)",
          "Signal Processing (eess.SP)"
        ]
      },
      {
        "id": "arXiv:2212.01728",
        "title": "An ISAC-based Beam Alignment Approach for Enhancing Terahertz Network  Coverage",
        "abstract": "Terahertz (THz) communication can provide high-capacity data transmissions benefiting from the ultra-broad bandwidth, but suffers from severe propagation loss and line-of-sight (LoS) blockage that limits the network coverage. To compensate for the loss, narrow beams are required. But they in turn bring in THz narrow beam alignment challenges in mobile networks, resulting in poor link connection and severe latency that degrades the network performance. This paper exploits integrated sensing and communication (ISAC) technology to reduce the THz beam misalignment caused by LoS blockage and user mobility. In line with the 5G-specified beam management, we propose a joint reference signal (RS) and the synchronization signal block (SSB)-based sensing scheme to predict the need for beam switches, and signalling in advance to prevent beam misalignment, named the joint SSB and RS-based sensing (JSRS) scheme. By reducing the impact of imperfect sensing, we further provide a time-frequency sensing resource allocation scheme that minimizes the beam misalignment probability. Considering the sensing trade-offs and THz features, we provide stochastic geometry-based expressions for coverage probability and network throughput, revealing design insights into ISAC-THz network deployment and resource allocation. In urban use cases tested, numerical results show that the JSRS scheme averagely reduces 80% probability of beam misalignment, and enhances the coverage probability by about 75%, compared to the network with 5G-required positioning ability. Despite the imperfect sensing accuracy, JSRS scheme can achieve comparable performance to that of the ideal-sensing case. By exploiting the angular perception ability of directional beams, JSRS scheme helps mitigate the challenge of narrow beam management to enhance the network coverage and throughput, showing the great potential of ISAC-aided THz communication. ",
        "url": "https://arxiv.org/abs/2212.01728",
        "authors": [
          "Wenrong Chen",
          "Lingxiang Li",
          "Zhi Chen",
          "Boyu Ning",
          "Guangjian Wang",
          "Tony Quek"
        ],
        "subjectives": [
          "Information Theory (cs.IT)"
        ]
      },
      {
        "id": "arXiv:2212.01735",
        "title": "Neural Fourier Filter Bank",
        "abstract": "We present a novel method to provide efficient and highly detailed reconstructions. Inspired by wavelets, our main idea is to learn a neural field that decompose the signal both spatially and frequency-wise. We follow the recent grid-based paradigm for spatial decomposition, but unlike existing work, encourage specific frequencies to be stored in each grid via Fourier features encodings. We then apply a multi-layer perceptron with sine activations, taking these Fourier encoded features in at appropriate layers so that higher-frequency components are accumulated on top of lower-frequency components sequentially, which we sum up to form the final output. We demonstrate that our method outperforms the state of the art regarding model compactness and efficiency on multiple tasks: 2D image fitting, 3D shape reconstruction, and neural radiance fields. ",
        "url": "https://arxiv.org/abs/2212.01735",
        "authors": [
          "Zhijie Wu",
          "Yuhe Jin",
          "Kwang Moo Yi"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Artificial Intelligence (cs.AI)",
          "Graphics (cs.GR)"
        ]
      },
      {
        "id": "arXiv:2212.01736",
        "title": "Downlink Transmission with Heterogeneous URLLC Services: Discrete  Signaling With Single-User Decoding",
        "abstract": "The problem of designing downlink transmission schemes for supporting heterogeneous ultra-reliable low-latency communications (URLLC) and/or with other types of services is investigated. We consider the broadcast channel, where the base station sends superimposed signals to multiple users. Under heterogeneous blocklength constraints, strong users who are URLLC users cannot wait to receive the entire transmission frame and perform successive interference cancellation (SIC) due to stringent latency requirements, in contrast to the conventional infinite blocklength cases. Even if SIC is feasible, SIC may be imperfect under finite blocklength constraints. To cope with the heterogeneity in latency and reliability requirements, we propose a practical downlink transmission scheme with discrete signaling and single-user decoding (SUD), i.e., without SIC. We carefully design the discrete input distributions to enable efficient SUD by exploiting the structural interference. Furthermore, we derive the second-order achievable rate under heterogenous blocklength and error probability constraints and use it to guide the design of channel coding and modulations. It is shown that in terms of achievable rate under short blocklength, the proposed scheme with regular quadrature amplitude modulations and SUD can operate extremely close to the benchmark schemes that assume perfect SIC with Gaussian signaling. ",
        "url": "https://arxiv.org/abs/2212.01736",
        "authors": [
          "Min Qiu",
          "Yu-Chih Huang",
          "Jinhong Yuan"
        ],
        "subjectives": [
          "Information Theory (cs.IT)",
          "Signal Processing (eess.SP)"
        ]
      },
      {
        "id": "arXiv:2212.01742",
        "title": "Lightweight Facial Attractiveness Prediction Using Dual Label  Distribution",
        "abstract": "Facial attractiveness prediction (FAP) aims to assess the facial attractiveness automatically based on human aesthetic perception. Previous methods using deep convolutional neural networks have boosted the performance, but their giant models lead to a deficiency in flexibility. Besides, most of them fail to take full advantage of the dataset. In this paper, we present a novel end-to-end FAP approach integrating dual label distribution and lightweight design. To make the best use of the dataset, the manual ratings, attractiveness score, and standard deviation are aggregated explicitly to construct a dual label distribution, including the attractiveness distribution and the rating distribution. Such distributions, as well as the attractiveness score, are optimized under a joint learning framework based on the label distribution learning (LDL) paradigm. As for the lightweight design, the data processing is simplified to minimum, and MobileNetV2 is selected as our backbone. Extensive experiments are conducted on two benchmark datasets, where our approach achieves promising results and succeeds in striking a balance between performance and efficiency. Ablation studies demonstrate that our delicately designed learning modules are indispensable and correlated. Additionally, the visualization indicates that our approach is capable of perceiving facial attractiveness and capturing attractive facial regions to facilitate semantic predictions. ",
        "url": "https://arxiv.org/abs/2212.01742",
        "authors": [
          "Shu Liu",
          "Enquan Huang",
          "Yan Xu",
          "Kexuan Wang",
          "Xiaoyan Kui",
          "Tao Lei",
          "Hongying Meng"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.01744",
        "title": "Statistical Physics of Deep Neural Networks: Initialization toward  Optimal Channels",
        "abstract": "In deep learning, neural networks serve as noisy channels between input data and its representation. This perspective naturally relates deep learning with the pursuit of constructing channels with optimal performance in information transmission and representation. While considerable efforts are concentrated on realizing optimal channel properties during network optimization, we study a frequently overlooked possibility that neural networks can be initialized toward optimal channels. Our theory, consistent with experimental validation, identifies primary mechanics underlying this unknown possibility and suggests intrinsic connections between statistical physics and deep learning. Unlike the conventional theories that characterize neural networks applying the classic mean-filed approximation, we offer analytic proof that this extensively applied simplification scheme is not valid in studying neural networks as information channels. To fill this gap, we develop a corrected mean-field framework applicable for characterizing the limiting behaviors of information propagation in neural networks without strong assumptions on inputs. Based on it, we propose an analytic theory to prove that mutual information maximization is realized between inputs and propagated signals when neural networks are initialized at dynamic isometry, a case where information transmits via norm-preserving mappings. These theoretical predictions are validated by experiments on real neural networks, suggesting the robustness of our theory against finite-size effects. Finally, we analyze our findings with information bottleneck theory to confirm the precise relations among dynamic isometry, mutual information maximization, and optimal channel properties in deep learning. ",
        "url": "https://arxiv.org/abs/2212.01744",
        "authors": [
          "Kangyu Weng",
          "Aohua Cheng",
          "Ziyang Zhang",
          "Pei Sun",
          "Yang Tian"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Disordered Systems and Neural Networks (cond-mat.dis-nn)",
          "Statistical Mechanics (cond-mat.stat-mech)",
          "Machine Learning (stat.ML)"
        ]
      },
      {
        "id": "arXiv:2212.01749",
        "title": "Semantic Graph Neural Network with Multi-measure Learning for  Semi-supervised Classification",
        "abstract": "Graph Neural Networks (GNNs) have attracted increasing attention in recent years and have achieved excellent performance in semi-supervised node classification tasks. The success of most GNNs relies on one fundamental assumption, i.e., the original graph structure data is available. However, recent studies have shown that GNNs are vulnerable to the complex underlying structure of the graph, making it necessary to learn comprehensive and robust graph structures for downstream tasks, rather than relying only on the raw graph structure. In light of this, we seek to learn optimal graph structures for downstream tasks and propose a novel framework for semi-supervised classification. Specifically, based on the structural context information of graph and node representations, we encode the complex interactions in semantics and generate semantic graphs to preserve the global structure. Moreover, we develop a novel multi-measure attention layer to optimize the similarity rather than prescribing it a priori, so that the similarity can be adaptively evaluated by integrating measures. These graphs are fused and optimized together with GNN towards semi-supervised classification objective. Extensive experiments and ablation studies on six real-world datasets clearly demonstrate the effectiveness of our proposed model and the contribution of each component. ",
        "url": "https://arxiv.org/abs/2212.01749",
        "authors": [
          "Junchao Lin",
          "Yuan Wan",
          "Jingwen Xu",
          "Xingchen Qi"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.01758",
        "title": "Improving Zero-shot Generalization and Robustness of Multi-modal Models",
        "abstract": "Multi-modal image-text models such as CLIP and LiT have demonstrated impressive performance on image classification benchmarks and their zero-shot generalization ability is particularly exciting. While the top-5 zero-shot accuracies of these models are very high, the top-1 accuracies are much lower (over 25% gap in some cases). We investigate the reasons for this performance gap and find that many of the failure cases are caused by ambiguity in the text prompts. First, we develop a simple and efficient zero-shot post-hoc method to identify images whose top-1 prediction is likely to be incorrect, by measuring consistency of the predictions w.r.t. multiple prompts and image transformations. We show that our procedure better predicts mistakes, outperforming the popular max logit baseline on selective prediction tasks. Next, we propose a simple and efficient way to improve accuracy on such uncertain images by making use of the WordNet hierarchy; specifically we augment the original class by incorporating its parent and children from the semantic label hierarchy, and plug the augmentation into text promts. We conduct experiments on both CLIP and LiT models with five different ImageNet-based datasets. For CLIP, our method improves the top-1 accuracy by 17.13% on the uncertain subset and 3.6% on the entire ImageNet validation set. We also show that our method improves across ImageNet shifted datasets and other model architectures such as LiT. Our proposed method is hyperparameter-free, requires no additional model training and can be easily scaled to other large multi-modal architectures. ",
        "url": "https://arxiv.org/abs/2212.01758",
        "authors": [
          "Yunhao Ge",
          "Jie Ren",
          "Yuxiao Wang",
          "Andrew Gallagher",
          "Ming-Hsuan Yang",
          "Laurent Itti",
          "Hartwig Adam",
          "Balaji Lakshminarayanan",
          "Jiaping Zhao"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.01762",
        "title": "Self-supervised AutoFlow",
        "abstract": "Recently, AutoFlow has shown promising results on learning a training set for optical flow, but requires ground truth labels in the target domain to compute its search metric. Observing a strong correlation between the ground truth search metric and self-supervised losses, we introduce self-supervised AutoFlow to handle real-world videos without ground truth labels. Using self-supervised loss as the search metric, our self-supervised AutoFlow performs on par with AutoFlow on Sintel and KITTI where ground truth is available, and performs better on the real-world DAVIS dataset. We further explore using self-supervised AutoFlow in the (semi-)supervised setting and obtain competitive results against the state of the art. ",
        "url": "https://arxiv.org/abs/2212.01762",
        "authors": [
          "Hsin-Ping Huang",
          "Charles Herrmann",
          "Junhwa Hur",
          "Erika Lu",
          "Kyle Sargent",
          "Austin Stone",
          "Ming-Hsuan Yang",
          "Deqing Sun"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.01764",
        "title": "Synthesize Boundaries: A Boundary-aware Self-consistent Framework for  Weakly Supervised Salient Object Detection",
        "abstract": "Fully supervised salient object detection (SOD) has made considerable progress based on expensive and time-consuming data with pixel-wise annotations. Recently, to relieve the labeling burden while maintaining performance, some scribble-based SOD methods have been proposed. However, learning precise boundary details from scribble annotations that lack edge information is still difficult. In this paper, we propose to learn precise boundaries from our designed synthetic images and labels without introducing any extra auxiliary data. The synthetic image creates boundary information by inserting synthetic concave regions that simulate the real concave regions of salient objects. Furthermore, we propose a novel self-consistent framework that consists of a global integral branch (GIB) and a boundary-aware branch (BAB) to train a saliency detector. GIB aims to identify integral salient objects, whose input is the original image. BAB aims to help predict accurate boundaries, whose input is the synthetic image. These two branches are connected through a self-consistent loss to guide the saliency detector to predict precise boundaries while identifying salient objects. Experimental results on five benchmarks demonstrate that our method outperforms the state-of-the-art weakly supervised SOD methods and further narrows the gap with the fully supervised methods. ",
        "url": "https://arxiv.org/abs/2212.01764",
        "authors": [
          "Binwei Xu",
          "Haoran Liang",
          "Ronghua Liang",
          "Peng Chen"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.01767",
        "title": "ConfounderGAN: Protecting Image Data Privacy with Causal Confounder",
        "abstract": "The success of deep learning is partly attributed to the availability of massive data downloaded freely from the Internet. However, it also means that users' private data may be collected by commercial organizations without consent and used to train their models. Therefore, it's important and necessary to develop a method or tool to prevent unauthorized data exploitation. In this paper, we propose ConfounderGAN, a generative adversarial network (GAN) that can make personal image data unlearnable to protect the data privacy of its owners. Specifically, the noise produced by the generator for each image has the confounder property. It can build spurious correlations between images and labels, so that the model cannot learn the correct mapping from images to labels in this noise-added dataset. Meanwhile, the discriminator is used to ensure that the generated noise is small and imperceptible, thereby remaining the normal utility of the encrypted image for humans. The experiments are conducted in six image classification datasets, consisting of three natural object datasets and three medical datasets. The results demonstrate that our method not only outperforms state-of-the-art methods in standard settings, but can also be applied to fast encryption scenarios. Moreover, we show a series of transferability and stability experiments to further illustrate the effectiveness and superiority of our method. ",
        "url": "https://arxiv.org/abs/2212.01767",
        "authors": [
          "Qi Tian",
          "Kun Kuang",
          "Kelu Jiang",
          "Furui Liu",
          "Zhihua Wang",
          "Fei Wu"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Artificial Intelligence (cs.AI)"
        ]
      },
      {
        "id": "arXiv:2212.01768",
        "title": "3D Object Aided Self-Supervised Monocular Depth Estimation",
        "abstract": "Monocular depth estimation has been actively studied in fields such as robot vision, autonomous driving, and 3D scene understanding. Given a sequence of color images, unsupervised learning methods based on the framework of Structure-From-Motion (SfM) simultaneously predict depth and camera relative pose. However, dynamically moving objects in the scene violate the static world assumption, resulting in inaccurate depths of dynamic objects. In this work, we propose a new method to address such dynamic object movements through monocular 3D object detection. Specifically, we first detect 3D objects in the images and build the per-pixel correspondence of the dynamic pixels with the detected object pose while leaving the static pixels corresponding to the rigid background to be modeled with camera motion. In this way, the depth of every pixel can be learned via a meaningful geometry model. Besides, objects are detected as cuboids with absolute scale, which is used to eliminate the scale ambiguity problem inherent in monocular vision. Experiments on the KITTI depth dataset show that our method achieves State-of-The-Art performance for depth estimation. Furthermore, joint training of depth, camera motion and object pose also improves monocular 3D object detection performance. To the best of our knowledge, this is the first work that allows a monocular 3D object detection network to be fine-tuned in a self-supervised manner. ",
        "url": "https://arxiv.org/abs/2212.01768",
        "authors": [
          "Songlin Wei",
          "Guodong Chen",
          "Wenzheng Chi",
          "Zhenhua Wang",
          "Lining Sun"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.01770",
        "title": "Distributionally Robust Day-ahead Scheduling for Power-traffic Network  under a Potential Game Framework",
        "abstract": "Widespread utilization of electric vehicles (EVs) incurs more uncertainties and impacts on the scheduling of the power-transportation coupled network. This paper investigates optimal power scheduling for a power-transportation coupled network in the day-ahead energy market considering multiple uncertainties related to photovoltaic (PV) generation and the traffic demand of vehicles. The crux of this problem is to model the coupling relation between the two networks in the day-ahead scheduling stage and consider the intra-day spatial uncertainties of the source and load. Meanwhile, the flexible load with a certain adjustment margin is introduced to ensure the balance of supply and demand of power nodes and consume the renewable energy better. Furthermore, we show the interactions between the power system and EV users from a potential game-theoretic perspective, where the uncertainties are characterized by an ambiguity set. In order to ensure the individual optimality of the two networks in a unified framework in day-ahead power scheduling, a two-stage distributionally robust centralized optimization model is established to carry out the equilibrium of power-transportation coupled network. On this basis, a combination of the duality theory and the Benders decomposition is developed to solve the distributionally robust optimization (DRO) model. Simulations demonstrate that the proposed approach can obtain individual optimal and less conservative strategies. ",
        "url": "https://arxiv.org/abs/2212.01770",
        "authors": [
          "Haoran Deng",
          "Bo Yang",
          "Chao Ning",
          "Cailian Chen",
          "Xinping Guan"
        ],
        "subjectives": [
          "Systems and Control (eess.SY)"
        ]
      },
      {
        "id": "arXiv:2212.01771",
        "title": "Can Evolutionary Clustering Have Theoretical Guarantees?",
        "abstract": "Clustering is a fundamental problem in many areas, which aims to partition a given data set into groups based on some distance measure, such that the data points in the same group are similar while that in different groups are dissimilar. Due to its importance and NP-hardness, a lot of methods have been proposed, among which evolutionary algorithms are a class of popular ones. Evolutionary clustering has found many successful applications, but all the results are empirical, lacking theoretical support. This paper fills this gap by proving that the approximation performance of the GSEMO (a simple multi-objective evolutionary algorithm) for solving the three popular formulations of clustering, i.e., $k$-center, $k$-median and $k$-means, can be theoretically guaranteed. Furthermore, we prove that evolutionary clustering can have theoretical guarantees even when considering fairness, which tries to avoid algorithmic bias, and has recently been an important research topic in machine learning. ",
        "url": "https://arxiv.org/abs/2212.01771",
        "authors": [
          "Chao Qian"
        ],
        "subjectives": [
          "Neural and Evolutionary Computing (cs.NE)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.01790",
        "title": "Kernel Inversed Pyramidal Resizing Network for Efficient Pavement  Distress Recognition",
        "abstract": "Pavement Distress Recognition (PDR) is an important step in pavement inspection and can be powered by image-based automation to expedite the process and reduce labor costs. Pavement images are often in high-resolution with a low ratio of distressed to non-distressed areas. Advanced approaches leverage these properties via dividing images into patches and explore discriminative features in the scale space. However, these approaches usually suffer from information loss during image resizing and low efficiency due to complex learning frameworks. In this paper, we propose a novel and efficient method for PDR. A light network named the Kernel Inversed Pyramidal Resizing Network (KIPRN) is introduced for image resizing, and can be flexibly plugged into the image classification network as a pre-network to exploit resolution and scale information. In KIPRN, pyramidal convolution and kernel inversed convolution are specifically designed to mine discriminative information across different feature granularities and scales. The mined information is passed along to the resized images to yield an informative image pyramid to assist the image classification network for PDR. We applied our method to three well-known Convolutional Neural Networks (CNNs), and conducted an evaluation on a large-scale pavement image dataset named CQU-BPDD. Extensive results demonstrate that KIPRN can generally improve the pavement distress recognition of these CNN models and show that the simple combination of KIPRN and EfficientNet-B3 significantly outperforms the state-of-the-art patch-based method in both performance and efficiency. ",
        "url": "https://arxiv.org/abs/2212.01790",
        "authors": [
          "Rong Qin",
          "Luwen Huangfu",
          "Devon Hood",
          "James Ma",
          "Sheng Huang"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Artificial Intelligence (cs.AI)"
        ]
      },
      {
        "id": "arXiv:2212.01793",
        "title": "Hyperbolic Curvature Graph Neural Network",
        "abstract": "Hyperbolic space is emerging as a promising learning space for representation learning, owning to its exponential growth volume. Compared with the flat Euclidean space, the curved hyperbolic space is far more ambient and embeddable, particularly for datasets with implicit tree-like architectures, such as hierarchies and power-law distributions. On the other hand, the structure of a real-world network is usually intricate, with some regions being tree-like, some being flat, and others being circular. Directly embedding heterogeneous structural networks into a homogeneous embedding space unavoidably brings inductive biases and distortions. Inspiringly, the discrete curvature can well describe the local structure of a node and its surroundings, which motivates us to investigate the information conveyed by the network topology explicitly in improving geometric learning. To this end, we explore the properties of the local discrete curvature of graph topology and the continuous global curvature of embedding space. Besides, a Hyperbolic Curvature-aware Graph Neural Network, HCGNN, is further proposed. In particular, HCGNN utilizes the discrete curvature to lead message passing of the surroundings and adaptively adjust the continuous curvature simultaneously. Extensive experiments on node classification and link prediction tasks show that the proposed method outperforms various competitive models by a large margin in both high and low hyperbolic graph data. Case studies further illustrate the efficacy of discrete curvature in finding local clusters and alleviating the distortion caused by hyperbolic geometry. ",
        "url": "https://arxiv.org/abs/2212.01793",
        "authors": [
          "Menglin Yang",
          "Min Zhou",
          "Lujia Pan",
          "Irwin King"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Artificial Intelligence (cs.AI)"
        ]
      },
      {
        "id": "arXiv:2212.01806",
        "title": "Recognizing Object by Components with Human Prior Knowledge Enhances  Adversarial Robustness of Deep Neural Networks",
        "abstract": "Adversarial attacks can easily fool object recognition systems based on deep neural networks (DNNs). Although many defense methods have been proposed in recent years, most of them can still be adaptively evaded. One reason for the weak adversarial robustness may be that DNNs are only supervised by category labels and do not have part-based inductive bias like the recognition process of humans. Inspired by a well-known theory in cognitive psychology -- recognition-by-components, we propose a novel object recognition model ROCK (Recognizing Object by Components with human prior Knowledge). It first segments parts of objects from images, then scores part segmentation results with predefined human prior knowledge, and finally outputs prediction based on the scores. The first stage of ROCK corresponds to the process of decomposing objects into parts in human vision. The second stage corresponds to the decision process of the human brain. ROCK shows better robustness than classical recognition models across various attack settings. These results encourage researchers to rethink the rationality of currently widely-used DNN-based object recognition models and explore the potential of part-based models, once important but recently ignored, for improving robustness. ",
        "url": "https://arxiv.org/abs/2212.01806",
        "authors": [
          "Xiao Li",
          "Ziqi Wang",
          "Bo Zhang",
          "Fuchun Sun",
          "Xiaolin Hu"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Artificial Intelligence (cs.AI)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.01819",
        "title": "Towards Precise Flood Prediction via Hierachical Terrain Attention and  Multi-Scale Rainfall Guidance",
        "abstract": "With the deterioration of climate, the phenomenon of rain-induced flooding has become frequent. To mitigate its impact, recent works adopt convolutional neural networks or other variants to predict the floods. However, these methods directly force the model to reconstruct the raw pixels of water depth maps through constraining pixel-level differences, ignoring the high-level information contained in terrain features and rainfall patterns. To address this, we present a novel GAN-based framework for precise flood prediction, which incorporates hierarchical terrain spatial attention to help the model focus on spatially-salient areas of terrain features and constructs multi-scale rainfall embedding to extensively integrate rainfall pattern information into generation. To better adapt the model in various rainfall conditions, we leverage a rainfall regression loss for both the generator and the discriminator as additional supervision. Extensive evaluations on real catchment datasets demonstrate the superior performance of our method, which greatly surpasses the previous arts under different rainfall conditions. ",
        "url": "https://arxiv.org/abs/2212.01819",
        "authors": [
          "Feifei Wang",
          "Yong Wang",
          "Shaoqing Chen",
          "Bing Li",
          "Qidong Huang"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.01833",
        "title": "Understanding Sinusoidal Neural Networks",
        "abstract": "In this work, we investigate the representation capacity of multilayer perceptron networks that use the sine as activation function - sinusoidal neural networks. We show that the layer composition in such networks compacts information. For this, we prove that the composition of sinusoidal layers expands as a sum of sines consisting of a large number of new frequencies given by linear combinations of the weights of the network's first layer. We provide the expression of the corresponding amplitudes in terms of the Bessel functions and give an upper bound for them that can be used to control the resulting approximation. ",
        "url": "https://arxiv.org/abs/2212.01833",
        "authors": [
          "Tiago Novello"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.01842",
        "title": "GraphGDP: Generative Diffusion Processes for Permutation Invariant Graph  Generation",
        "abstract": "Graph generative models have broad applications in biology, chemistry and social science. However, modelling and understanding the generative process of graphs is challenging due to the discrete and high-dimensional nature of graphs, as well as permutation invariance to node orderings in underlying graph distributions. Current leading autoregressive models fail to capture the permutation invariance nature of graphs for the reliance on generation ordering and have high time complexity. Here, we propose a continuous-time generative diffusion process for permutation invariant graph generation to mitigate these issues. Specifically, we first construct a forward diffusion process defined by a stochastic differential equation (SDE), which smoothly converts graphs within the complex distribution to random graphs that follow a known edge probability. Solving the corresponding reverse-time SDE, graphs can be generated from newly sampled random graphs. To facilitate the reverse-time SDE, we newly design a position-enhanced graph score network, capturing the evolving structure and position information from perturbed graphs for permutation equivariant score estimation. Under the evaluation of comprehensive metrics, our proposed generative diffusion process achieves competitive performance in graph distribution learning. Experimental results also show that GraphGDP can generate high-quality graphs in only 24 function evaluations, much faster than previous autoregressive models. ",
        "url": "https://arxiv.org/abs/2212.01842",
        "authors": [
          "Han Huang",
          "Leilei Sun",
          "Bowen Du",
          "Yanjie Fu",
          "Weifeng Lv"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Social and Information Networks (cs.SI)"
        ]
      },
      {
        "id": "arXiv:2212.01844",
        "title": "Pair-Based Joint Encoding with Relational Graph Convolutional Networks  for Emotion-Cause Pair Extraction",
        "abstract": "Emotion-cause pair extraction (ECPE) aims to extract emotion clauses and corresponding cause clauses, which have recently received growing attention. Previous methods sequentially encode features with a specified order. They first encode the emotion and cause features for clause extraction and then combine them for pair extraction. This lead to an imbalance in inter-task feature interaction where features extracted later have no direct contact with the former. To address this issue, we propose a novel Pair-Based Joint Encoding (PBJE) network, which generates pairs and clauses features simultaneously in a joint feature encoding manner to model the causal relationship in clauses. PBJE can balance the information flow among emotion clauses, cause clauses and pairs. From a multi-relational perspective, we construct a heterogeneous undirected graph and apply the Relational Graph Convolutional Network (RGCN) to capture the various relationship between clauses and the relationship between pairs and clauses. Experimental results show that PBJE achieves state-of-the-art performance on the Chinese benchmark corpus. ",
        "url": "https://arxiv.org/abs/2212.01844",
        "authors": [
          "Junlong Liu",
          "Xichen Shang",
          "Qianli Ma"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)",
          "Artificial Intelligence (cs.AI)"
        ]
      },
      {
        "id": "arXiv:2212.01855",
        "title": "Pairing-Friendly Elliptic Curves: Revisited Taxonomy, Attacks and  Security Concern",
        "abstract": "Major families of pairing-friendly elliptic curves, including BN, BLS12, BLS24, KSS16, and KSS18 have recently been vulnerable to number field sieve (NFS) attacks. Due to the recent attacks on discrete logs in F_(q^k ), selecting such curves became relevant again. This paper revisited the topic of selecting pairing-friendly curves at different security levels. First, we expanded the classification given by Freeman et al. [1] by identifying new families that were not previously mentioned, such as a complete family with variable differentiation and new sparse families of curves. We discussed individual curves and a comprehensive framework for constructing parametric families. We estimated the security and assessed families of the pairing-friendly curve to discover families of curves better than BN, KSS, and BLS in terms of the required key size. We also evaluated the complexity of the optimal ate pairing that has never been discussed before, except by Barbulescu et al. [2]. We demonstrated that the recent attack (TNFS) on pairing needs to increase the key size. We compared families of curves in the context of key size and selected a suitable alternative to an elliptic curve. ",
        "url": "https://arxiv.org/abs/2212.01855",
        "authors": [
          "Mahender Kumar",
          "Satish Chand"
        ],
        "subjectives": [
          "Cryptography and Security (cs.CR)"
        ]
      },
      {
        "id": "arXiv:2212.01880",
        "title": "Deep Learning for Multiscale Damage Analysis via Physics-Informed  Recurrent Neural Network",
        "abstract": "Direct numerical simulation of hierarchical materials via homogenization-based concurrent multiscale models poses critical challenges for 3D large scale engineering applications, as the computation of highly nonlinear and path-dependent material constitutive responses at the lower scale causes prohibitively high computational costs. In this work, we propose a physics-informed data-driven deep learning model as an efficient surrogate to emulate the effective responses of heterogeneous microstructures under irreversible elasto-plastic hardening and softening deformation. Our contribution contains several major innovations. First, we propose a novel training scheme to generate arbitrary loading sequences in the sampling space confined by deformation constraints where the simulation cost of homogenizing microstructural responses per sequence is dramatically reduced via mechanistic reduced-order models. Second, we develop a new sequential learner that incorporates thermodynamics consistent physics constraints by customizing training loss function and data flow architecture. We additionally demonstrate the integration of trained surrogate within the framework of classic multiscale finite element solver. Our numerical experiments indicate that our model shows a significant accuracy improvement over pure data-driven emulator and a dramatic efficiency boost than reduced models. We believe our data-driven model provides a computationally efficient and mechanics consistent alternative for classic constitutive laws beneficial for potential high-throughput simulations that needs material homogenization of irreversible behaviors. ",
        "url": "https://arxiv.org/abs/2212.01880",
        "authors": [
          "Shiguang Deng"
        ],
        "subjectives": [
          "Computational Engineering, Finance, and Science (cs.CE)"
        ]
      },
      {
        "id": "arXiv:2212.01882",
        "title": "Visualizing Contributor Code Competency for PyPI Libraries: Preliminary  Results",
        "abstract": "Python is known to be used by beginners to professional programmers. Python provides functionality to its community of users through PyPI libraries, which allows developers to reuse functionalities to an application. However, it is unknown the extent to which these PyPI libraries require proficient code in their implementation. We conjecture that PyPI contributors may decide to implement more advanced Pythonic code, or stick with more basic Python code. Are complex codes only committed by few contributors, or only to specific files? The new idea in this paper is to confirm who and where complex code is implemented. Hence, we present a visualization to show the relationship between proficient code, contributors, and files. Analyzing four PyPI projects, we are able to explore which files contain more elegant code, and which contributors committed to these files. Our results show that most files contain more basic competency files, and that not every contributor contributes competent code. We show how~our visualization is able to summarize such information, and opens up different possibilities for understanding how to make elegant contributions. ",
        "url": "https://arxiv.org/abs/2212.01882",
        "authors": [
          "Indira Febriyanti",
          "Raula Gaikovina Kula",
          "Ruksit Rojpaisarnkit",
          "Kanchanok Kannee",
          "Yusuf Sulistyo Nugroho",
          "Kenichi Matsumoto"
        ],
        "subjectives": [
          "Software Engineering (cs.SE)"
        ]
      },
      {
        "id": "arXiv:2212.01893",
        "title": "Joint Self-Supervised Image-Volume Representation Learning with  Intra-Inter Contrastive Clustering",
        "abstract": "Collecting large-scale medical datasets with fully annotated samples for training of deep networks is prohibitively expensive, especially for 3D volume data. Recent breakthroughs in self-supervised learning (SSL) offer the ability to overcome the lack of labeled training samples by learning feature representations from unlabeled data. However, most current SSL techniques in the medical field have been designed for either 2D images or 3D volumes. In practice, this restricts the capability to fully leverage unlabeled data from numerous sources, which may include both 2D and 3D data. Additionally, the use of these pre-trained networks is constrained to downstream tasks with compatible data dimensions. In this paper, we propose a novel framework for unsupervised joint learning on 2D and 3D data modalities. Given a set of 2D images or 2D slices extracted from 3D volumes, we construct an SSL task based on a 2D contrastive clustering problem for distinct classes. The 3D volumes are exploited by computing vectored embedding at each slice and then assembling a holistic feature through deformable self-attention mechanisms in Transformer, allowing incorporating long-range dependencies between slices inside 3D volumes. These holistic features are further utilized to define a novel 3D clustering agreement-based SSL task and masking embedding prediction inspired by pre-trained language models. Experiments on downstream tasks, such as 3D brain segmentation, lung nodule detection, 3D heart structures segmentation, and abnormal chest X-ray detection, demonstrate the effectiveness of our joint 2D and 3D SSL approach. We improve plain 2D Deep-ClusterV2 and SwAV by a significant margin and also surpass various modern 2D and 3D SSL approaches. ",
        "url": "https://arxiv.org/abs/2212.01893",
        "authors": [
          "Duy M. H. Nguyen",
          "Hoang Nguyen",
          "Mai T. N. Truong",
          "Tri Cao",
          "Binh T. Nguyen",
          "Nhat Ho",
          "Paul Swoboda",
          "Shadi Albarqouni",
          "Pengtao Xie",
          "Daniel Sonntag"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.01896",
        "title": "A proactive autoscaling and energy-efficient VM allocation framework  using online multi-resource neural network for cloud data center",
        "abstract": "This work proposes an energy-efficient resource provisioning and allocation framework to meet the dynamic demands of future applications. The frequent variations in a cloud user's resource demand lead 'to the problem of excess power consumption, resource wastage, performance, and Quality-of-Service degradation. The proposed framework addresses these challenges by matching the application's predicted resource requirement with the resource capacity of VMs precisely and thereby consolidating the entire load on the minimum number of energy-efficient physical machines. The three consecutive contributions of the proposed work are: Online Multi-Resource Feed-forward Neural Network to forecast the multiple resource demands concurrently for future applications; autoscaling of VMs based on the clustering of the predicted resource requirements; allocation of the scaled VMs on the energy-efficient PMs. The integrated approach successively optimizes resource utilization, saves energy and automatically adapts to the changes in future application resource demand. The proposed framework is evaluated by using real workload traces of the benchmark Google Cluster Dataset and compared against different scenarios including energy-efficient VM placement with resource prediction only, VMP without resource prediction and autoscaling, and optimal VMP with autoscaling based on actual resource utilization. The observed results demonstrate that the proposed integrated approach achieves near-optimal performance against optimal VMP and outperforms rest of the VMPs in terms of power saving and resource utilization up to 88.5% and 21.12% respectively. In addition, the OM-FNN predictor shows better accuracy, lesser time and space complexity over a traditional single-input and single-output feed-forward neural network predictor. ",
        "url": "https://arxiv.org/abs/2212.01896",
        "authors": [
          "Deepika Saxena",
          "Ashutosh Kumar Singh"
        ],
        "subjectives": [
          "Distributed, Parallel, and Cluster Computing (cs.DC)"
        ]
      },
      {
        "id": "arXiv:2212.01904",
        "title": "Graph Representation Learning for Wireless Communications",
        "abstract": "Wireless networks are inherently graph-structured, which can be utilized in graph representation learning to solve complex wireless network optimization problems. In graph representation learning, feature vectors for each entity in the network are calculated such that they capture spatial and temporal dependencies in their local and global neighbourhoods. Graph neural networks (GNNs) are powerful tools to solve these complex problems because of their expressive representation and reasoning power. In this paper, the potential of graph representation learning and GNNs in wireless networks is presented. An overview of graph learning is provided which covers the fundamentals and concepts such as feature design over graphs, GNNs, and their design principles. Potential of graph representation learning in wireless networks is presented via few exemplary use cases and some initial results on the GNN-based access point selection for cell-free massive MIMO systems. ",
        "url": "https://arxiv.org/abs/2212.01904",
        "authors": [
          "Maryam Mohsenivatani",
          "Samad Ali",
          "Vismika Ranasinghe",
          "Nandana Rajatheva",
          "Matti Latva-Aho"
        ],
        "subjectives": [
          "Information Theory (cs.IT)"
        ]
      },
      {
        "id": "arXiv:2212.01906",
        "title": "Combining multiple matchers for fingerprint verification: A case study  in biosecure network of excellence",
        "abstract": "We report on experiments for the fingerprint modality conducted during the First BioSecure Residential Workshop. Two reference systems for fingerprint verification have been tested together with two additional non-reference systems. These systems follow different approaches of fingerprint processing and are discussed in detail. Fusion experiments I volving different combinations of the available systems are presented. The experimental results show that the best recognition strategy involves both minutiae-based and correlation-based measurements. Regarding the fusion experiments, the best relative improvement is obtained when fusing systems that are based on heterogeneous strategies for feature extraction and/or matching. The best combinations of two/three/four systems always include the best individual systems whereas the best verification performance is obtained when combining all the available systems. ",
        "url": "https://arxiv.org/abs/2212.01906",
        "authors": [
          "Fernando Alonso-Fernandez",
          "Julian Fierrez-Aguilar",
          "Hartwig Fronthaler",
          "Klaus Kollreider",
          "Javier Ortega-Garcia",
          "Joaquin Gonzalez-Rodriguez",
          "Josef Bigun"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.01923",
        "title": "Query-Driven Knowledge Base Completion using Multimodal Path Fusion over  Multimodal Knowledge Graph",
        "abstract": "Over the past few years, large knowledge bases have been constructed to store massive amounts of knowledge. However, these knowledge bases are highly incomplete, for example, over 70% of people in Freebase have no known place of birth. To solve this problem, we propose a query-driven knowledge base completion system with multimodal fusion of unstructured and structured information. To effectively fuse unstructured information from the Web and structured information in knowledge bases to achieve good performance, our system builds multimodal knowledge graphs based on question answering and rule inference. We propose a multimodal path fusion algorithm to rank candidate answers based on different paths in the multimodal knowledge graphs, achieving much better performance than question answering, rule inference and a baseline fusion algorithm. To improve system efficiency, query-driven techniques are utilized to reduce the runtime of our system, providing fast responses to user queries. Extensive experiments have been conducted to demonstrate the effectiveness and efficiency of our system. ",
        "url": "https://arxiv.org/abs/2212.01923",
        "authors": [
          "Yang Peng",
          "Daisy Zhe Wang"
        ],
        "subjectives": [
          "Databases (cs.DB)",
          "Artificial Intelligence (cs.AI)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.01927",
        "title": "Label Encoding for Regression Networks",
        "abstract": "Deep neural networks are used for a wide range of regression problems. However, there exists a significant gap in accuracy between specialized approaches and generic direct regression in which a network is trained by minimizing the squared or absolute error of output labels. Prior work has shown that solving a regression problem with a set of binary classifiers can improve accuracy by utilizing well-studied binary classification algorithms. We introduce binary-encoded labels (BEL), which generalizes the application of binary classification to regression by providing a framework for considering arbitrary multi-bit values when encoding target values. We identify desirable properties of suitable encoding and decoding functions used for the conversion between real-valued and binary-encoded labels based on theoretical and empirical study. These properties highlight a tradeoff between classification error probability and error-correction capabilities of label encodings. BEL can be combined with off-the-shelf task-specific feature extractors and trained end-to-end. We propose a series of sample encoding, decoding, and training loss functions for BEL and demonstrate they result in lower error than direct regression and specialized approaches while being suitable for a diverse set of regression problems, network architectures, and evaluation metrics. BEL achieves state-of-the-art accuracies for several regression benchmarks. Code is available at https://github.com/ubc-aamodt-group/BEL_regression. ",
        "url": "https://arxiv.org/abs/2212.01927",
        "authors": [
          "Deval Shah",
          "Zi Yu Xue",
          "Tor M. Aamodt"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.01939",
        "title": "Winning the CityLearn Challenge: Adaptive Optimization with Evolutionary  Search under Trajectory-based Guidance",
        "abstract": "Modern power systems will have to face difficult challenges in the years to come: frequent blackouts in urban areas caused by high power demand peaks, grid instability exacerbated by intermittent renewable generation, and global climate change amplified by rising carbon emissions. While current practices are growingly inadequate, the path to widespread adoption of artificial intelligence (AI) methods is hindered by missing aspects of trustworthiness. The CityLearn Challenge is an exemplary opportunity for researchers from multiple disciplines to investigate the potential of AI to tackle these pressing issues in the energy domain, collectively modeled as a reinforcement learning (RL) task. Multiple real-world challenges faced by contemporary RL techniques are embodied in the problem formulation. In this paper, we present a novel method using the solution function of optimization as policies to compute actions for sequential decision-making, while notably adapting the parameters of the optimization model from online observations. Algorithmically, this is achieved by an evolutionary algorithm under a novel trajectory-based guidance scheme. Formally, the global convergence property is established. Our agent ranked first in the latest 2021 CityLearn Challenge, being able to achieve superior performance in almost all metrics while maintaining some key aspects of interpretability. ",
        "url": "https://arxiv.org/abs/2212.01939",
        "authors": [
          "Vanshaj Khattar",
          "Ming Jin"
        ],
        "subjectives": [
          "Systems and Control (eess.SY)",
          "Machine Learning (cs.LG)",
          "Neural and Evolutionary Computing (cs.NE)"
        ]
      },
      {
        "id": "arXiv:2212.01941",
        "title": "Systematic Design and Evaluation of Social Determinants of Health  Ontology (SDoHO)",
        "abstract": "Social determinants of health (SDoH) have a significant impact on health outcomes and well-being. Addressing SDoH is the key to reducing healthcare inequalities and transforming a sick care system into a health-promoting system. To address the SDOH terminology gap and better embed relevant elements in advanced biomedical informatics, we propose an SDoH ontology (SDoHO), which represents fundamental SDoH factors and their relationships in a standardized and measurable way. The ontology formally models classes, relationships, and constraints based on multiple SDoH-related resources. Expert review and coverage evaluation, using clinical notes data and a national survey, showed satisfactory results. SDoHO could potentially play an essential role in providing a foundation for a comprehensive understanding of the associations between SDoH and health outcomes and providing a path toward health equity across populations. ",
        "url": "https://arxiv.org/abs/2212.01941",
        "authors": [
          "Yifang Dang",
          "Fang Li",
          "Xinyue Hu",
          "Vipina K. Keloth",
          "Meng Zhang",
          "Sunyang Fu",
          "Jingcheng Du",
          "J. Wilfred Fan",
          "Muhammad F. Amith",
          "Evan Yu",
          "Hongfang Liu",
          "Xiaoqian Jiang",
          "Hua Xu",
          "Cui Tao"
        ],
        "subjectives": [
          "Computers and Society (cs.CY)"
        ]
      },
      {
        "id": "arXiv:2212.01944",
        "title": "Learning Automata-Based Task Knowledge Representation from Large-Scale  Generative Language Models",
        "abstract": "Automata-based representations play an important role in control and planning in sequential decision-making, but obtaining high-level task knowledge for building automata is often difficult. Although large-scale generative language models (GLMs) can help automatically distill task knowledge, the textual outputs from GLMs are not directly utilizable in sequential decision-making. We resolve this problem by proposing a novel algorithm named GLM2FSA, which obtains high-level task knowledge, represented in a finite state automaton (FSA), from a given brief description of the task goal. GLM2FSA sends queries to a GLM for task knowledge in textual form and then builds a FSA to represent the textual knowledge. This algorithm fills the gap between text and automata-based representations, and the constructed FSA can be directly utilized in sequential decision-making. We provide examples to demonstrate how GLM2FSA constructs FSAs to represent knowledge encoded in the texts generated by the large-scale GLMs. ",
        "url": "https://arxiv.org/abs/2212.01944",
        "authors": [
          "Yunhao Yang",
          "Jean-Rapha\u00ebl Gaglione",
          "Ufuk Topcu"
        ],
        "subjectives": [
          "Formal Languages and Automata Theory (cs.FL)",
          "Computation and Language (cs.CL)"
        ]
      },
      {
        "id": "arXiv:2212.01957",
        "title": "CSTAR: Towards Compact and STructured Deep Neural Networks with  Adversarial Robustness",
        "abstract": "Model compression and model defense for deep neural networks (DNNs) have been extensively and individually studied. Considering the co-importance of model compactness and robustness in practical applications, several prior works have explored to improve the adversarial robustness of the sparse neural networks. However, the structured sparse models obtained by the exiting works suffer severe performance degradation for both benign and robust accuracy, thereby causing a challenging dilemma between robustness and structuredness of the compact DNNs. To address this problem, in this paper, we propose CSTAR, an efficient solution that can simultaneously impose the low-rankness-based Compactness, high STructuredness and high Adversarial Robustness on the target DNN models. By formulating the low-rankness and robustness requirement within the same framework and globally determining the ranks, the compressed DNNs can simultaneously achieve high compression performance and strong adversarial robustness. Evaluations for various DNN models on different datasets demonstrate the effectiveness of CSTAR. Compared with the state-of-the-art robust structured pruning methods, CSTAR shows consistently better performance. For instance, when compressing ResNet-18 on CIFAR-10, CSTAR can achieve up to 20.07% and 11.91% improvement for benign accuracy and robust accuracy, respectively. For compressing ResNet-18 with 16x compression ratio on Imagenet, CSTAR can obtain 8.58% benign accuracy gain and 4.27% robust accuracy gain compared to the existing robust structured pruning method. ",
        "url": "https://arxiv.org/abs/2212.01957",
        "authors": [
          "Huy Phan",
          "Miao Yin",
          "Yang Sui",
          "Bo Yuan",
          "Saman Zonouz"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.01959",
        "title": "INGeo: Accelerating Instant Neural Scene Reconstruction with Noisy  Geometry Priors",
        "abstract": "We present a method that accelerates reconstruction of 3D scenes and objects, aiming to enable instant reconstruction on edge devices such as mobile phones and AR/VR headsets. While recent works have accelerated scene reconstruction training to minute/second-level on high-end GPUs, there is still a large gap to the goal of instant training on edge devices which is yet highly desired in many emerging applications such as immersive AR/VR. To this end, this work aims to further accelerate training by leveraging geometry priors of the target scene. Our method proposes strategies to alleviate the noise of the imperfect geometry priors to accelerate the training speed on top of the highly optimized Instant-NGP. On the NeRF Synthetic dataset, our work uses half of the training iterations to reach an average test PSNR of >30. ",
        "url": "https://arxiv.org/abs/2212.01959",
        "authors": [
          "Chaojian Li",
          "Bichen Wu",
          "Albert Pumarola",
          "Peizhao Zhang",
          "Yingyan Lin",
          "Peter Vajda"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.01968",
        "title": "Dissimilar Nodes Improve Graph Active Learning",
        "abstract": "Training labels for graph embedding algorithms could be costly to obtain in many practical scenarios. Active learning (AL) algorithms are very helpful to obtain the most useful labels for training while keeping the total number of label queries under a certain budget. The existing Active Graph Embedding framework proposes to use centrality score, density score, and entropy score to evaluate the value of unlabeled nodes, and it has been shown to be capable of bringing some improvement to the node classification tasks of Graph Convolutional Networks. However, when evaluating the importance of unlabeled nodes, it fails to consider the influence of existing labeled nodes on the value of unlabeled nodes. In other words, given the same unlabeled node, the computed informative score is always the same and is agnostic to the labeled node set. With the aim to address this limitation, in this work, we introduce 3 dissimilarity-based information scores for active learning: feature dissimilarity score (FDS), structure dissimilarity score (SDS), and embedding dissimilarity score (EDS). We find out that those three scores are able to take the influence of the labeled set on the value of unlabeled candidates into consideration, boosting our AL performance. According to experiments, our newly proposed scores boost the classification accuracy by 2.1% on average and are capable of generalizing to different Graph Neural Network architectures. ",
        "url": "https://arxiv.org/abs/2212.01968",
        "authors": [
          "Zhicheng Ren",
          "Yifu Yuan",
          "Yuxin Wu",
          "Xiaxuan Gao",
          "Yewen Wang",
          "Yizhou Sun"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Artificial Intelligence (cs.AI)"
        ]
      },
      {
        "id": "arXiv:2212.01976",
        "title": "FedCC: Robust Federated Learning against Model Poisoning Attacks",
        "abstract": "Federated Learning has emerged to cope with raising concerns about privacy breaches in using Machine or Deep Learning models. This new paradigm allows the leverage of deep learning models in a distributed manner, enhancing privacy preservation. However, the server's blindness to local datasets introduces its vulnerability to model poisoning attacks and data heterogeneity, tampering with the global model performance. Numerous works have proposed robust aggregation algorithms and defensive mechanisms, but the approaches are orthogonal to individual attacks or issues. FedCC, the proposed method, provides robust aggregation by comparing the Centered Kernel Alignment of Penultimate Layers Representations. The experiment results on FedCC demonstrate that it mitigates untargeted and targeted model poisoning or backdoor attacks while also being effective in non-Independently and Identically Distributed data environments. By applying FedCC against untargeted attacks, global model accuracy is recovered the most. Against targeted backdoor attacks, FedCC nullified attack confidence while preserving the test accuracy. Most of the experiment results outstand the baseline methods. ",
        "url": "https://arxiv.org/abs/2212.01976",
        "authors": [
          "Hyejun Jeong",
          "Hamin Son",
          "Seohu Lee",
          "Jayun Hyun",
          "Tai-Myoung Chung"
        ],
        "subjectives": [
          "Cryptography and Security (cs.CR)",
          "Artificial Intelligence (cs.AI)"
        ]
      },
      {
        "id": "arXiv:2212.01985",
        "title": "ObjectMatch: Robust Registration using Canonical Object Correspondences",
        "abstract": "We present ObjectMatch, a semantic and object-centric camera pose estimation for RGB-D SLAM pipelines. Modern camera pose estimators rely on direct correspondences of overlapping regions between frames; however, they cannot align camera frames with little or no overlap. In this work, we propose to leverage indirect correspondences obtained via semantic object identification. For instance, when an object is seen from the front in one frame and from the back in another frame, we can provide additional pose constraints through canonical object correspondences. We first propose a neural network to predict such correspondences on a per-pixel level, which we then combine in our energy formulation with state-of-the-art keypoint matching solved with a joint Gauss-Newton optimization. In a pairwise setting, our method improves registration recall of state-of-the-art feature matching from 77% to 87% overall and from 21% to 52% in pairs with 10% or less inter-frame overlap. In registering RGB-D sequences, our method outperforms cutting-edge SLAM baselines in challenging, low frame-rate scenarios, achieving more than 35% reduction in trajectory error in multiple scenes. ",
        "url": "https://arxiv.org/abs/2212.01985",
        "authors": [
          "Can G\u00fcmeli",
          "Angela Dai",
          "Matthias Nie\u00dfner"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.01992",
        "title": "Fast and accurate factorized neural transducer for text adaption of  end-to-end speech recognition models",
        "abstract": "Neural transducer is now the most popular end-to-end model for speech recognition, due to its naturally streaming ability. However, it is challenging to adapt it with text-only data. Factorized neural transducer (FNT) model was proposed to mitigate this problem. The improved adaptation ability of FNT on text-only adaptation data came at the cost of lowered accuracy compared to the standard neural transducer model. We propose several methods to improve the performance of the FNT model. They are: adding CTC criterion during training, adding KL divergence loss during adaptation, using a pre-trained language model to seed the vocabulary predictor, and an efficient adaptation approach by interpolating the vocabulary predictor with the n-gram language model. A combination of these approaches results in a relative word-error-rate reduction of 9.48\\% from the standard FNT model. Furthermore, n-gram interpolation with the vocabulary predictor improves the adaptation speed hugely with satisfactory adaptation performance. ",
        "url": "https://arxiv.org/abs/2212.01992",
        "authors": [
          "Rui Zhao",
          "Jian Xue",
          "Partha Parthasarathy",
          "Veljko Miljanic",
          "Jinyu Li"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)",
          "Sound (cs.SD)",
          "Audio and Speech Processing (eess.AS)"
        ]
      },
      {
        "id": "arXiv:2212.01995",
        "title": "AOP-Miner: Approximate Order-Preserving Pattern Mining for Time Series",
        "abstract": "The order-preserving pattern mining can be regarded as discovering frequent trends in time series, since the same order-preserving pattern has the same relative order which can represent a trend. However, in the case where data noise is present, the relative orders of many meaningful patterns are usually similar rather than the same. To mine similar relative orders in time series, this paper addresses an approximate order-preserving pattern (AOP) mining method based on (delta-gamma) distance to effectively measure the similarity, and proposes an algorithm called AOP-Miner to mine AOPs according to global and local approximation parameters. AOP-Miner adopts a pattern fusion strategy to generate candidate patterns generation and employs the screening strategy to calculate the supports of candidate patterns. Experimental results validate that AOP-Miner outperforms other competitive methods and can find more similar trends in time series. ",
        "url": "https://arxiv.org/abs/2212.01995",
        "authors": [
          "Yan Li",
          "Jin Liu",
          "Yingchun Guo",
          "Jing Liu",
          "Youxi Wu"
        ],
        "subjectives": [
          "Databases (cs.DB)"
        ]
      },
      {
        "id": "arXiv:2212.02000",
        "title": "Wish I Can Feel What You Feel: A Neural Approach for Empathetic Response  Generation",
        "abstract": "Expressing empathy is important in everyday conversations, and exploring how empathy arises is crucial in automatic response generation. Most previous approaches consider only a single factor that affects empathy. However, in practice, empathy generation and expression is a very complex and dynamic psychological process. A listener needs to find out events which cause a speaker's emotions (emotion cause extraction), project the events into some experience (knowledge extension), and express empathy in the most appropriate way (communication mechanism). To this end, we propose a novel approach, which integrates the three components - emotion cause, knowledge graph, and communication mechanism for empathetic response generation. Experimental results on the benchmark dataset demonstrate the effectiveness of our method and show that incorporating the key components generates more informative and empathetic responses. ",
        "url": "https://arxiv.org/abs/2212.02000",
        "authors": [
          "Yangbin Chen",
          "Chunfeng Liang"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)",
          "Human-Computer Interaction (cs.HC)"
        ]
      },
      {
        "id": "arXiv:2212.02003",
        "title": "Bayesian Learning with Information Gain Provably Bounds Risk for a  Robust Adversarial Defense",
        "abstract": "We present a new algorithm to learn a deep neural network model robust against adversarial attacks. Previous algorithms demonstrate an adversarially trained Bayesian Neural Network (BNN) provides improved robustness. We recognize the adversarial learning approach for approximating the multi-modal posterior distribution of a Bayesian model can lead to mode collapse; consequently, the model's achievements in robustness and performance are sub-optimal. Instead, we first propose preventing mode collapse to better approximate the multi-modal posterior distribution. Second, based on the intuition that a robust model should ignore perturbations and only consider the informative content of the input, we conceptualize and formulate an information gain objective to measure and force the information learned from both benign and adversarial training instances to be similar. Importantly. we prove and demonstrate that minimizing the information gain objective allows the adversarial risk to approach the conventional empirical risk. We believe our efforts provide a step toward a basis for a principled method of adversarially training BNNs. Our model demonstrate significantly improved robustness--up to 20%--compared with adversarial training and Adv-BNN under PGD attacks with 0.035 distortion on both CIFAR-10 and STL-10 datasets. ",
        "url": "https://arxiv.org/abs/2212.02003",
        "authors": [
          "Bao Gia Doan",
          "Ehsan Abbasnejad",
          "Javen Qinfeng Shi",
          "Damith C. Ranasinghe"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Cryptography and Security (cs.CR)",
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.02006",
        "title": "HierarchyFL: Heterogeneous Federated Learning via Hierarchical  Self-Distillation",
        "abstract": "Federated learning (FL) has been recognized as a privacy-preserving distributed machine learning paradigm that enables knowledge sharing among various heterogeneous artificial intelligence (AIoT) devices through centralized global model aggregation. FL suffers from model inaccuracy and slow convergence due to the model heterogeneity of the AIoT devices involved. Although various existing methods try to solve the bottleneck of the model heterogeneity problem, most of them improve the accuracy of heterogeneous models in a coarse-grained manner, which makes it still a great challenge to deploy large-scale AIoT devices. To alleviate the negative impact of this problem and take full advantage of the diversity of each heterogeneous model, we propose an efficient framework named HierarchyFL, which uses a small amount of public data for efficient and scalable knowledge across a variety of differently structured models. By using self-distillation and our proposed ensemble library, each hierarchical model can intelligently learn from each other on cloud servers. Experimental results on various well-known datasets show that HierarchyFL can not only maximize the knowledge sharing among various heterogeneous models in large-scale AIoT systems, but also greatly improve the model performance of each involved heterogeneous AIoT device. ",
        "url": "https://arxiv.org/abs/2212.02006",
        "authors": [
          "Jun Xia",
          "Yi Zhang",
          "Zhihao Yue",
          "Ming Hu",
          "Xian Wei",
          "Mingsong Chen"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.02010",
        "title": "Multi Agent Path Finding using Evolutionary Game Theory",
        "abstract": "In this paper, we consider the problem of path finding for a set of homogeneous and autonomous agents navigating a previously unknown stochastic environment. In our problem setting, each agent attempts to maximize a given utility function while respecting safety properties. Our solution is based on ideas from evolutionary game theory, namely replicating policies that perform well and diminishing ones that do not. We do a comprehensive comparison with related multiagent planning methods, and show that our technique beats state of the art RL algorithms in minimizing path length by nearly 30% in large spaces. We show that our algorithm is computationally faster than deep RL methods by at least an order of magnitude. We also show that it scales better with an increase in the number of agents as compared to other methods, path planning methods in particular. Lastly, we empirically prove that the policies that we learn are evolutionarily stable and thus impervious to invasion by any other policy. ",
        "url": "https://arxiv.org/abs/2212.02010",
        "authors": [
          "Sheryl Paul",
          "Jyotirmoy V. Deshmukh"
        ],
        "subjectives": [
          "Multiagent Systems (cs.MA)",
          "Artificial Intelligence (cs.AI)",
          "Computer Science and Game Theory (cs.GT)"
        ]
      },
      {
        "id": "arXiv:2212.02014",
        "title": "Med-Query: Steerable Parsing of 9-DoF Medical Anatomies with Query  Embedding",
        "abstract": "Automatic parsing of human anatomies at instance-level from 3D computed tomography (CT) scans is a prerequisite step for many clinical applications. The presence of pathologies, broken structures or limited field-of-view (FOV) all can make anatomy parsing algorithms vulnerable. In this work, we explore how to exploit and conduct the prosperous detection-then-segmentation paradigm in 3D medical data, and propose a steerable, robust, and efficient computing framework for detection, identification, and segmentation of anatomies in CT scans. Considering complicated shapes, sizes and orientations of anatomies, without lose of generality, we present the nine degrees-of-freedom (9-DoF) pose estimation solution in full 3D space using a novel single-stage, non-hierarchical forward representation. Our whole framework is executed in a steerable manner where any anatomy of interest can be directly retrieved to further boost the inference efficiency. We have validated the proposed method on three medical imaging parsing tasks of ribs, spine, and abdominal organs. For rib parsing, CT scans have been annotated at the rib instance-level for quantitative evaluation, similarly for spine vertebrae and abdominal organs. Extensive experiments on 9-DoF box detection and rib instance segmentation demonstrate the effectiveness of our framework (with the identification rate of 97.0% and the segmentation Dice score of 90.9%) in high efficiency, compared favorably against several strong baselines (e.g., CenterNet, FCOS, and nnU-Net). For spine identification and segmentation, our method achieves a new state-of-the-art result on the public CTSpine1K dataset. Last, we report highly competitive results in multi-organ segmentation at FLARE22 competition. Our annotations, code and models will be made publicly available at: https://github.com/alibaba-damo-academy/Med_Query. ",
        "url": "https://arxiv.org/abs/2212.02014",
        "authors": [
          "Heng Guo",
          "Jianfeng Zhang",
          "Ke Yan",
          "Le Lu",
          "Minfeng Xu"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.02017",
        "title": "GNN-SL: Sequence Labeling Based on Nearest Examples via GNN",
        "abstract": "To better handle long-tail cases in the sequence labeling (SL) task, in this work, we introduce graph neural networks sequence labeling (GNN-SL), which augments the vanilla SL model output with similar tagging examples retrieved from the whole training set. Since not all the retrieved tagging examples benefit the model prediction, we construct a heterogeneous graph, and leverage graph neural networks (GNNs) to transfer information between the retrieved tagging examples and the input word sequence. The augmented node which aggregates information from neighbors is used to do prediction. This strategy enables the model to directly acquire similar tagging examples and improves the general quality of predictions. We conduct a variety of experiments on three typical sequence labeling tasks: Named Entity Recognition (NER), Part of Speech Tagging (POS), and Chinese Word Segmentation (CWS) to show the significant performance of our GNN-SL. Notably, GNN-SL achieves SOTA results of 96.9 (+0.2) on PKU, 98.3 (+0.4) on CITYU, 98.5 (+0.2) on MSR, and 96.9 (+0.2) on AS for the CWS task, and results comparable to SOTA performances on NER datasets, and POS datasets. ",
        "url": "https://arxiv.org/abs/2212.02017",
        "authors": [
          "Shuhe Wang",
          "Yuxian Meng",
          "Rongbin Ouyang",
          "Jiwei Li",
          "Tianwei Zhang",
          "Lingjuan Lyu",
          "Guoyin Wang"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)"
        ]
      },
      {
        "id": "arXiv:2212.02031",
        "title": "Prototypical Residual Networks for Anomaly Detection and Localization",
        "abstract": "Anomaly detection and localization are widely used in industrial manufacturing for its efficiency and effectiveness. Anomalies are rare and hard to collect and supervised models easily over-fit to these seen anomalies with a handful of abnormal samples, producing unsatisfactory performance. On the other hand, anomalies are typically subtle, hard to discern, and of various appearance, making it difficult to detect anomalies and let alone locate anomalous regions. To address these issues, we propose a framework called Prototypical Residual Network (PRN), which learns feature residuals of varying scales and sizes between anomalous and normal patterns to accurately reconstruct the segmentation maps of anomalous regions. PRN mainly consists of two parts: multi-scale prototypes that explicitly represent the residual features of anomalies to normal patterns; a multisize self-attention mechanism that enables variable-sized anomalous feature learning. Besides, we present a variety of anomaly generation strategies that consider both seen and unseen appearance variance to enlarge and diversify anomalies. Extensive experiments on the challenging and widely used MVTec AD benchmark show that PRN outperforms current state-of-the-art unsupervised and supervised methods. We further report SOTA results on three additional datasets to demonstrate the effectiveness and generalizability of PRN. ",
        "url": "https://arxiv.org/abs/2212.02031",
        "authors": [
          "Hui Zhang",
          "Zuxuan Wu",
          "Zheng Wang",
          "Zhineng Chen",
          "Yu-Gang Jiang"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.02042",
        "title": "Refiner: Data Refining against Gradient Leakage Attacks in Federated  Learning",
        "abstract": "Federated Learning (FL) is pervasive in privacy-focused IoT environments since it enables avoiding privacy leakage by training models with gradients instead of data. Recent works show the uploaded gradients can be employed to reconstruct data, i.e., gradient leakage attacks, and several defenses are designed to alleviate the risk by tweaking the gradients. However, these defenses exhibit weak resilience against threatening attacks, as the effectiveness builds upon the unrealistic assumptions that deep neural networks are simplified as linear models. In this paper, without such unrealistic assumptions, we present a novel defense, called Refiner, instead of perturbing gradients, which refines ground-truth data to craft robust data that yields sufficient utility but with the least amount of privacy information, and then the gradients of robust data are uploaded. To craft robust data, Refiner promotes the gradients of critical parameters associated with robust data to close ground-truth ones while leaving the gradients of trivial parameters to safeguard privacy. Moreover, to exploit the gradients of trivial parameters, Refiner utilizes a well-designed evaluation network to steer robust data far away from ground-truth data, thereby alleviating privacy leakage risk. Extensive experiments across multiple benchmark datasets demonstrate the superior defense effectiveness of Refiner at defending against state-of-the-art threats. ",
        "url": "https://arxiv.org/abs/2212.02042",
        "authors": [
          "Mingyuan Fan",
          "Cen Chen",
          "Chengyu Wang",
          "Wenmeng Zhou",
          "Jun Huang",
          "Ximeng Liu",
          "Wenzhong Guo"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Cryptography and Security (cs.CR)"
        ]
      },
      {
        "id": "arXiv:2212.02046",
        "title": "Algorithm and Hardware Co-Design of Energy-Efficient LSTM Networks for  Video Recognition with Hierarchical Tucker Tensor Decomposition",
        "abstract": "Long short-term memory (LSTM) is a type of powerful deep neural network that has been widely used in many sequence analysis and modeling applications. However, the large model size problem of LSTM networks make their practical deployment still very challenging, especially for the video recognition tasks that require high-dimensional input data. Aiming to overcome this limitation and fully unlock the potentials of LSTM models, in this paper we propose to perform algorithm and hardware co-design towards high-performance energy-efficient LSTM networks. At algorithm level, we propose to develop fully decomposed hierarchical Tucker (FDHT) structure-based LSTM, namely FDHT-LSTM, which enjoys ultra-low model complexity while still achieving high accuracy. In order to fully reap such attractive algorithmic benefit, we further develop the corresponding customized hardware architecture to support the efficient execution of the proposed FDHT-LSTM model. With the delicate design of memory access scheme, the complicated matrix transformation can be efficiently supported by the underlying hardware without any access conflict in an on-the-fly way. Our evaluation results show that both the proposed ultra-compact FDHT-LSTM models and the corresponding hardware accelerator achieve very high performance. Compared with the state-of-the-art compressed LSTM models, FDHT-LSTM enjoys both order-of-magnitude reduction in model size and significant accuracy improvement across different video recognition datasets. Meanwhile, compared with the state-of-the-art tensor decomposed model-oriented hardware TIE, our proposed FDHT-LSTM architecture achieves better performance in throughput, area efficiency and energy efficiency, respectively on LSTM-Youtube workload. For LSTM-UCF workload, our proposed design also outperforms TIE with higher throughput, higher energy efficiency and comparable area efficiency. ",
        "url": "https://arxiv.org/abs/2212.02046",
        "authors": [
          "Yu Gong",
          "Miao Yin",
          "Lingyi Huang",
          "Chunhua Deng",
          "Yang Sui",
          "Bo Yuan"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.02048",
        "title": "Hodge Decomposition of the Remittance Network on the XRP ledger in the  Price Hike of January 2018",
        "abstract": "This study analyzes the remittance transaction recorded on the XRP ledger for ETH and USD from July 2017 to Jun 2018, including the bubble period in early 2018. Using the Hodge decomposition, we estimate the ``loop flow'' in the international remittance of cryptoassets during the bubble period. We found characteristic differences between those fiat currencies and cryptoassets during the bubble period. For ETH, there was a significant increase in the loop flow during the cryptoasset price peak. This might be related to money laundering or arbitrage transaction. There was a slight increase in the loop flow for USD during the cryptoasset price peak. ",
        "url": "https://arxiv.org/abs/2212.02048",
        "authors": [
          "Yuichi Ikeda",
          "Abhijit Chakraborty"
        ],
        "subjectives": [
          "Cryptography and Security (cs.CR)"
        ]
      },
      {
        "id": "arXiv:2212.02055",
        "title": "Graph Convolutional Neural Networks with Diverse Negative Samples via  Decomposed Determinant Point Processes",
        "abstract": "Graph convolutional networks (GCNs) have achieved great success in graph representation learning by extracting high-level features from nodes and their topology. Since GCNs generally follow a message-passing mechanism, each node aggregates information from its first-order neighbour to update its representation. As a result, the representations of nodes with edges between them should be positively correlated and thus can be considered positive samples. However, there are more non-neighbour nodes in the whole graph, which provide diverse and useful information for the representation update. Two non-adjacent nodes usually have different representations, which can be seen as negative samples. Besides the node representations, the structural information of the graph is also crucial for learning. In this paper, we used quality-diversity decomposition in determinant point processes (DPP) to obtain diverse negative samples. When defining a distribution on diverse subsets of all non-neighbouring nodes, we incorporate both graph structure information and node representations. Since the DPP sampling process requires matrix eigenvalue decomposition, we propose a new shortest-path-base method to improve computational efficiency. Finally, we incorporate the obtained negative samples into the graph convolution operation. The ideas are evaluated empirically in experiments on node classification tasks. These experiments show that the newly proposed methods not only improve the overall performance of standard representation learning but also significantly alleviate over-smoothing problems. ",
        "url": "https://arxiv.org/abs/2212.02055",
        "authors": [
          "Wei Duan",
          "Junyu Xuan",
          "Maoying Qiao",
          "Jie Lu"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.02057",
        "title": "DA-CIL: Towards Domain Adaptive Class-Incremental 3D Object Detection",
        "abstract": "Deep learning has achieved notable success in 3D object detection with the advent of large-scale point cloud datasets. However, severe performance degradation in the past trained classes, i.e., catastrophic forgetting, still remains a critical issue for real-world deployment when the number of classes is unknown or may vary. Moreover, existing 3D class-incremental detection methods are developed for the single-domain scenario, which fail when encountering domain shift caused by different datasets, varying environments, etc. In this paper, we identify the unexplored yet valuable scenario, i.e., class-incremental learning under domain shift, and propose a novel 3D domain adaptive class-incremental object detection framework, DA-CIL, in which we design a novel dual-domain copy-paste augmentation method to construct multiple augmented domains for diversifying training distributions, thereby facilitating gradual domain adaptation. Then, multi-level consistency is explored to facilitate dual-teacher knowledge distillation from different domains for domain adaptive class-incremental learning. Extensive experiments on various datasets demonstrate the effectiveness of the proposed method over baselines in the domain adaptive class-incremental learning scenario. ",
        "url": "https://arxiv.org/abs/2212.02057",
        "authors": [
          "Ziyuan Zhao",
          "Mingxi Xu",
          "Peisheng Qian",
          "Ramanpreet Singh Pahwa",
          "Richard Chang"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Artificial Intelligence (cs.AI)",
          "Image and Video Processing (eess.IV)"
        ]
      },
      {
        "id": "arXiv:2212.02060",
        "title": "Resilience Evaluation of Entropy Regularized Logistic Networks with  Probabilistic Cost",
        "abstract": "The demand for resilient logistics networks has increased because of recent disasters. When we consider optimization problems, entropy regularization is a powerful tool for the diversification of a solution. In this study, we proposed a method for designing a resilient logistics network based on entropy regularization. Moreover, we proposed a method for analytical resilience criteria to reduce the ambiguity of resilience. First, we modeled the logistics network, including factories, distribution bases, and sales outlets in an efficient framework using entropy regularization. Next, we formulated a resilience criterion based on probabilistic cost and Kullback--Leibler divergence. Finally, our method was performed using a simple logistics network, and the resilience of the three logistics plans designed by entropy regularization was demonstrated. ",
        "url": "https://arxiv.org/abs/2212.02060",
        "authors": [
          "Koshi Oishi",
          "Yota Hashizume",
          "Tomohiko Jimbo",
          "Hirotaka Kaji",
          "Kenji Kashima"
        ],
        "subjectives": [
          "Systems and Control (eess.SY)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.02075",
        "title": "Differentiated Federated Reinforcement Learning for Dynamic and  Heterogeneous Network",
        "abstract": "The modern dynamic and heterogeneous network brings differential environments with respective state transition probability to agents, which leads to the local strategy trap problem of traditional federated reinforcement learning (FRL) based network optimization algorithm. To solve this problem, we propose a novel Differentiated Federated Reinforcement Learning (DFRL), which evolves the global policy model integration and local inference with the global policy model in traditional FRL to a collaborative learning process with parallel global trends learning and differential local policy model learning. In the DFRL, the local policy learning model is adaptively updated with the global trends model and local environment and achieves better differentiated adaptation. We evaluate the outperformance of the proposal compared with the state-of-the-art FRL in a classical CartPole game with heterogeneous environments. Furthermore, we implement the proposal in the heterogeneous Space-air-ground Integrated Network (SAGIN) for the classical traffic offloading problem in network. The simulation result shows that the proposal shows better global performance and fairness than baselines in terms of throughput, delay, and packet drop rate. ",
        "url": "https://arxiv.org/abs/2212.02075",
        "authors": [
          "Fengxiao Tang",
          "Yilin Yang",
          "Xin Yao",
          "Ming Zhao",
          "Nei Kato"
        ],
        "subjectives": [
          "Networking and Internet Architecture (cs.NI)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.02077",
        "title": "DL-SLOT: Dynamic LiDAR SLAM and object tracking based on collaborative  graph optimization",
        "abstract": "Ego-pose estimation and dynamic object tracking are two critical problems for autonomous driving systems. The solutions to these problems are generally based on their respective assumptions, \\ie{the static world assumption for simultaneous localization and mapping (SLAM) and the accurate ego-pose assumption for object tracking}. However, these assumptions are challenging to hold in dynamic road scenarios, where SLAM and object tracking become closely correlated. Therefore, we propose DL-SLOT, a dynamic LiDAR SLAM and object tracking method, to simultaneously address these two coupled problems. This method integrates the state estimations of both the autonomous vehicle and the stationary and dynamic objects in the environment into a unified optimization framework. First, we used object detection to identify all points belonging to potentially dynamic objects. Subsequently, a LiDAR odometry was conducted using the filtered point cloud. Simultaneously, we proposed a sliding window-based object association method that accurately associates objects according to the historical trajectories of tracked objects. The ego-states and those of the stationary and dynamic objects are integrated into the sliding window-based collaborative graph optimization. The stationary objects are subsequently restored from the potentially dynamic object set. Finally, a global pose-graph is implemented to eliminate the accumulated error. Experiments on KITTI datasets demonstrate that our method achieves better accuracy than SLAM and object tracking baseline methods. This confirms that solving SLAM and object tracking simultaneously is mutually advantageous, dramatically improving the robustness and accuracy of SLAM and object tracking in dynamic road scenarios. ",
        "url": "https://arxiv.org/abs/2212.02077",
        "authors": [
          "Xuebo Tian",
          "Zhongyang Zhu",
          "Junqiao Zhao",
          "Gengxuan Tian",
          "Chen Ye"
        ],
        "subjectives": [
          "Robotics (cs.RO)"
        ]
      },
      {
        "id": "arXiv:2212.02081",
        "title": "YolOOD: Utilizing Object Detection Concepts for Out-of-Distribution  Detection",
        "abstract": "Out-of-distribution (OOD) detection has attracted a large amount of attention from the machine learning research community in recent years due to its importance in deployed systems. Most of the previous studies focused on the detection of OOD samples in the multi-class classification task. However, OOD detection in the multi-label classification task remains an underexplored domain. In this research, we propose YolOOD - a method that utilizes concepts from the object detection domain to perform OOD detection in the multi-label classification task. Object detection models have an inherent ability to distinguish between objects of interest (in-distribution) and irrelevant objects (e.g., OOD objects) on images that contain multiple objects from different categories. These abilities allow us to convert a regular object detection model into an image classifier with inherent OOD detection capabilities with just minor changes. We compare our approach to state-of-the-art OOD detection methods and demonstrate YolOOD's ability to outperform these methods on a comprehensive suite of in-distribution and OOD benchmark datasets. ",
        "url": "https://arxiv.org/abs/2212.02081",
        "authors": [
          "Alon Zolfi",
          "Guy Amit",
          "Amit Baras",
          "Satoru Koda",
          "Ikuya Morikawa",
          "Yuval Elovici",
          "Asaf Shabtai"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.02082",
        "title": "Hierarchical Contrast for Unsupervised Skeleton-based Action  Representation Learning",
        "abstract": "This paper targets unsupervised skeleton-based action representation learning and proposes a new Hierarchical Contrast (HiCo) framework. Different from the existing contrastive-based solutions that typically represent an input skeleton sequence into instance-level features and perform contrast holistically, our proposed HiCo represents the input into multiple-level features and performs contrast in a hierarchical manner. Specifically, given a human skeleton sequence, we represent it into multiple feature vectors of different granularities from both temporal and spatial domains via sequence-to-sequence (S2S) encoders and unified downsampling modules. Besides, the hierarchical contrast is conducted in terms of four levels: instance level, domain level, clip level, and part level. Moreover, HiCo is orthogonal to the S2S encoder, which allows us to flexibly embrace state-of-the-art S2S encoders. Extensive experiments on four datasets, i.e., NTU-60, NTU-120, PKU-MMD I and II, show that HiCo achieves a new state-of-the-art for unsupervised skeleton-based action representation learning in two downstream tasks including action recognition and retrieval, and its learned action representation is of good transferability. Besides, we also show that our framework is effective for semi-supervised skeleton-based action recognition. Our code is available at https://github.com/HuiGuanLab/HiCo. ",
        "url": "https://arxiv.org/abs/2212.02082",
        "authors": [
          "Jianfeng Dong",
          "Shengkai Sun",
          "Zhonglin Liu",
          "Shujie Chen",
          "Baolong Liu",
          "Xun Wang"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.02084",
        "title": "End-to-end Recording Device Identification Based on Deep Representation  Learning",
        "abstract": "Deep learning techniques have achieved specific results in recording device source identification. The recording device source features include spatial information and certain temporal information. However, most recording device source identification methods based on deep learning only use spatial representation learning from recording device source features, which cannot make full use of recording device source information. Therefore, in this paper, to fully explore the spatial information and temporal information of recording device source, we propose a new method for recording device source identification based on the fusion of spatial feature information and temporal feature information by using an end-to-end framework. From a feature perspective, we designed two kinds of networks to extract recording device source spatial and temporal information. Afterward, we use the attention mechanism to adaptively assign the weight of spatial information and temporal information to obtain fusion features. From a model perspective, our model uses an end-to-end framework to learn the deep representation from spatial feature and temporal feature and train using deep and shallow loss to joint optimize our network. This method is compared with our previous work and baseline system. The results show that the proposed method is better than our previous work and baseline system under general conditions. ",
        "url": "https://arxiv.org/abs/2212.02084",
        "authors": [
          "Chunyan Zeng",
          "Dongliang Zhu",
          "Zhifeng Wang",
          "Minghu Wu",
          "Wei Xiong",
          "Nan Zhao"
        ],
        "subjectives": [
          "Sound (cs.SD)",
          "Audio and Speech Processing (eess.AS)"
        ]
      },
      {
        "id": "arXiv:2212.02090",
        "title": "Breaking the Spurious Causality of Conditional Generation via Fairness  Intervention with Corrective Sampling",
        "abstract": "Trying to capture the sample-label relationship, conditional generative models often end up inheriting the spurious correlation in the training dataset, giving label-conditional distributions that are severely imbalanced in another latent attribute. To mitigate such undesirable correlations engraved into generative models, which we call spurious causality, we propose a general two-step strategy. (a) Fairness Intervention (FI): Emphasize the minority samples that are hard to be generated due to the spurious correlation in the training dataset. (b) Corrective Sampling (CS): Filter the generated samples explicitly to follow the desired label-conditional latent attribute distribution. We design the fairness intervention for various degrees of supervision on the spurious attribute, including unsupervised, weakly-supervised, and semi-supervised scenarios. Our experimental results show that the proposed FICS can successfully resolve the spurious correlation in generated samples on various datasets. ",
        "url": "https://arxiv.org/abs/2212.02090",
        "authors": [
          "Junhyun Nam",
          "Sangwoo Mo",
          "Jaeho Lee",
          "Jinwoo Shin"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Artificial Intelligence (cs.AI)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.02096",
        "title": "Inspired by Norbert Wiener: FeedBack Loop Network Learning Incremental  Knowledge for Driver Attention Prediction and Beyond",
        "abstract": "The problem of predicting driver attention from the driving perspective is gaining the increasing research focuses due to its remarkable significance for autonomous driving and assisted driving systems. Driving experience is extremely important for driver attention prediction, a skilled driver is able to effortlessly predict oncoming danger (before it becomes salient) based on driving experience and quickly pay attention on the corresponding zones. However, the nonobjective driving experience is difficult to model, so a mechanism simulating driver experience accumulation procedure is absent in existing methods, and the existing methods usually follow the technique line of saliency prediction methods to predict driver attention. In this paper, we propose a FeedBack Loop Network (FBLNet), which attempts to model the driving experience accumulation procedure. By over-and-over iterations, FBLNet generates the incremental knowledge that carries rich historically-accumulative long-term temporal information. The incremental knowledge to our model is like the driving experience to humans. Under the guidance of the incremental knowledge, our model fuses the CNN feature and Transformer feature that are extracted from the input image to predict driver attention. Our model exhibits solid advantage over existing methods, achieving an average 10.3% performance improvement on three public datasets. ",
        "url": "https://arxiv.org/abs/2212.02096",
        "authors": [
          "Yilong Chen",
          "Zhixiong Nan"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.02127",
        "title": "FaceQAN: Face Image Quality Assessment Through Adversarial Noise  Exploration",
        "abstract": "Recent state-of-the-art face recognition (FR) approaches have achieved impressive performance, yet unconstrained face recognition still represents an open problem. Face image quality assessment (FIQA) approaches aim to estimate the quality of the input samples that can help provide information on the confidence of the recognition decision and eventually lead to improved results in challenging scenarios. While much progress has been made in face image quality assessment in recent years, computing reliable quality scores for diverse facial images and FR models remains challenging. In this paper, we propose a novel approach to face image quality assessment, called FaceQAN, that is based on adversarial examples and relies on the analysis of adversarial noise which can be calculated with any FR model learned by using some form of gradient descent. As such, the proposed approach is the first to link image quality to adversarial attacks. Comprehensive (cross-model as well as model-specific) experiments are conducted with four benchmark datasets, i.e., LFW, CFP-FP, XQLFW and IJB-C, four FR models, i.e., CosFace, ArcFace, CurricularFace and ElasticFace, and in comparison to seven state-of-the-art FIQA methods to demonstrate the performance of FaceQAN. Experimental results show that FaceQAN achieves competitive results, while exhibiting several desirable characteristics. ",
        "url": "https://arxiv.org/abs/2212.02127",
        "authors": [
          "\u017diga Babnik",
          "Peter Peer",
          "Vitomir \u0160truc"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.02136",
        "title": "Adaptive Configuration for Heterogeneous Participants in Decentralized  Federated Learning",
        "abstract": "Data generated at the network edge can be processed locally by leveraging the paradigm of edge computing (EC). Aided by EC, decentralized federated learning (DFL), which overcomes the single-point-of-failure problem in the parameter server (PS) based federated learning, is becoming a practical and popular approach for machine learning over distributed data. However, DFL faces two critical challenges, \\ie, system heterogeneity and statistical heterogeneity introduced by edge devices. To ensure fast convergence with the existence of slow edge devices, we present an efficient DFL method, termed FedHP, which integrates adaptive control of both local updating frequency and network topology to better support the heterogeneous participants. We establish a theoretical relationship between local updating frequency and network topology regarding model training performance and obtain a convergence upper bound. Upon this, we propose an optimization algorithm, that adaptively determines local updating frequencies and constructs the network topology, so as to speed up convergence and improve the model accuracy. Evaluation results show that the proposed FedHP can reduce the completion time by about 51% and improve model accuracy by at least 5% in heterogeneous scenarios, compared with the baselines. ",
        "url": "https://arxiv.org/abs/2212.02136",
        "authors": [
          "Yunming Liao",
          "Yang Xu",
          "Hongli Xu",
          "Lun Wang",
          "Chen Qian"
        ],
        "subjectives": [
          "Networking and Internet Architecture (cs.NI)"
        ]
      },
      {
        "id": "arXiv:2212.02143",
        "title": "Impact of Domain-Adapted Multilingual Neural Machine Translation in the  Medical Domain",
        "abstract": "Multilingual Neural Machine Translation (MNMT) models leverage many language pairs during training to improve translation quality for low-resource languages by transferring knowledge from high-resource languages. We study the quality of a domain-adapted MNMT model in the medical domain for English-Romanian with automatic metrics and a human error typology annotation which includes terminology-specific error categories. We compare the out-of-domain MNMT with the in-domain adapted MNMT. The in-domain MNMT model outperforms the out-of-domain MNMT in all measured automatic metrics and produces fewer terminology errors. ",
        "url": "https://arxiv.org/abs/2212.02143",
        "authors": [
          "Miguel Rios",
          "Raluca-Maria Chereji",
          "Alina Secara",
          "Dragos Ciobanu"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)"
        ]
      },
      {
        "id": "arXiv:2212.02175",
        "title": "Momentum Decoding: Open-ended Text Generation As Graph Exploration",
        "abstract": "Open-ended text generation with autoregressive language models (LMs) is one of the core tasks in natural language processing. However, maximization-based decoding methods (e.g., greedy/beam search) often lead to the degeneration problem, i.e., the generated text is unnatural and contains undesirable repetitions. Existing solutions to this problem either introduce randomness prone to incoherence or require a look-ahead mechanism that demands extra computational overhead. In this study, we formulate open-ended text generation from a new perspective, i.e., we view it as an exploration process within a directed graph. Thereby, we understand the phenomenon of degeneration as circular loops within the directed graph. Based on our formulation, we propose a novel decoding method -- \\textit{momentum decoding} -- which encourages the LM to \\textit{greedily} explore new nodes outside the current graph. Meanwhile, it also allows the LM to return to the existing nodes with a momentum downgraded by a pre-defined resistance function. We extensively test our approach on three benchmarks from different domains through automatic and human evaluations. The results show that momentum decoding performs comparably with the current state of the art while enjoying notably improved inference speed and computation FLOPs. Furthermore, we conduct a detailed analysis to reveal the merits and inner workings of our approach. Our codes and other related resources are publicly available at https://github.com/gmftbyGMFTBY/MomentumDecoding. ",
        "url": "https://arxiv.org/abs/2212.02175",
        "authors": [
          "Tian Lan",
          "Yixuan Su",
          "Shuhang Liu",
          "Heyan Huang",
          "Xian-Ling Mao"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)"
        ]
      },
      {
        "id": "arXiv:2212.02181",
        "title": "Perceive, Interact, Predict: Learning Dynamic and Static Clues for  End-to-End Motion Prediction",
        "abstract": "Motion prediction is highly relevant to the perception of dynamic objects and static map elements in the scenarios of autonomous driving. In this work, we propose PIP, the first end-to-end Transformer-based framework which jointly and interactively performs online mapping, object detection and motion prediction. PIP leverages map queries, agent queries and mode queries to encode the instance-wise information of map elements, agents and motion intentions, respectively. Based on the unified query representation, a differentiable multi-task interaction scheme is proposed to exploit the correlation between perception and prediction. Even without human-annotated HD map or agent's historical tracking trajectory as guidance information, PIP realizes end-to-end multi-agent motion prediction and achieves better performance than tracking-based and HD-map-based methods. PIP provides comprehensive high-level information of the driving scene (vectorized static map and dynamic objects with motion information), and contributes to the downstream planning and control. Code and models will be released for facilitating further research. ",
        "url": "https://arxiv.org/abs/2212.02181",
        "authors": [
          "Bo Jiang",
          "Shaoyu Chen",
          "Xinggang Wang",
          "Bencheng Liao",
          "Tianheng Cheng",
          "Jiajie Chen",
          "Helong Zhou",
          "Qian Zhang",
          "Wenyu Liu",
          "Chang Huang"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Robotics (cs.RO)"
        ]
      },
      {
        "id": "arXiv:2212.02182",
        "title": "Anomaly Detection in Power Markets and Systems",
        "abstract": "The widespread use of information and communication technology (ICT) over the course of the last decades has been a primary catalyst behind the digitalization of power systems. Meanwhile, as the utilization rate of the Internet of Things (IoT) continues to rise along with recent advancements in ICT, the need for secure and computationally efficient monitoring of critical infrastructures like the electrical grid and the agents that participate in it is growing. A cyber-physical system, such as the electrical grid, may experience anomalies for a number of different reasons. These may include physical defects, mistakes in measurement and communication, cyberattacks, and other similar occurrences. The goal of this study is to emphasize what the most common incidents are with power systems and to give an overview and classification of the most common ways to find problems, starting with the consumer/prosumer end working up to the primary power producers. In addition, this article aimed to discuss the methods and techniques, such as artificial intelligence (AI) that are used to identify anomalies in the power systems and markets. ",
        "url": "https://arxiv.org/abs/2212.02182",
        "authors": [
          "Ugur Halden",
          "Umit Cali",
          "Ferhat Ozgur Catak",
          "Salvatore D'Arco",
          "Francisco Bilendo"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Signal Processing (eess.SP)"
        ]
      },
      {
        "id": "arXiv:2212.02191",
        "title": "Partial Variance Reduction improves Non-Convex Federated learning on  heterogeneous data",
        "abstract": "Data heterogeneity across clients is a key challenge in federated learning. Prior works address this by either aligning client and server models or using control variates to correct client model drift. Although these methods achieve fast convergence in convex or simple non-convex problems, the performance in over-parameterized models such as deep neural networks is lacking. In this paper, we first revisit the widely used FedAvg algorithm in a deep neural network to understand how data heterogeneity influences the gradient updates across the neural network layers. We observe that while the feature extraction layers are learned efficiently by FedAvg, the substantial diversity of the final classification layers across clients impedes the performance. Motivated by this, we propose to correct model drift by variance reduction only on the final layers. We demonstrate that this significantly outperforms existing benchmarks at a similar or lower communication cost. We furthermore provide proof for the convergence rate of our algorithm. ",
        "url": "https://arxiv.org/abs/2212.02191",
        "authors": [
          "Bo Li",
          "Mikkel N. Schmidt",
          "Tommy S. Alstr\u00f8m",
          "Sebastian U. Stich"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Distributed, Parallel, and Cluster Computing (cs.DC)"
        ]
      },
      {
        "id": "arXiv:2212.02199",
        "title": "Legal Prompt Engineering for Multilingual Legal Judgement Prediction",
        "abstract": "Legal Prompt Engineering (LPE) or Legal Prompting is a process to guide and assist a large language model (LLM) with performing a natural legal language processing (NLLP) skill. Our goal is to use LPE with LLMs over long legal documents for the Legal Judgement Prediction (LJP) task. We investigate the performance of zero-shot LPE for given facts in case-texts from the European Court of Human Rights (in English) and the Federal Supreme Court of Switzerland (in German, French and Italian). Our results show that zero-shot LPE is better compared to the baselines, but it still falls short compared to current state of the art supervised approaches. Nevertheless, the results are important, since there was 1) no explicit domain-specific data used - so we show that the transfer to the legal domain is possible for general-purpose LLMs, and 2) the LLMs where directly applied without any further training or fine-tuning - which in turn saves immensely in terms of additional computational costs. ",
        "url": "https://arxiv.org/abs/2212.02199",
        "authors": [
          "Dietrich Trautmann",
          "Alina Petrova",
          "Frank Schilder"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)",
          "Artificial Intelligence (cs.AI)"
        ]
      },
      {
        "id": "arXiv:2212.02230",
        "title": "A Hybrid Evolutionary Approach to Solve University Course Allocation  Problem",
        "abstract": "This paper discusses various types of constraints, difficulties and solutions to overcome the challenges regarding university course allocation problem. A hybrid evolutionary algorithm has been defined combining Local Repair Algorithm and Modified Genetic Algorithm to generate the best course assignment. After analyzing the collected dataset, all the necessary constraints were formulated. These constraints manage to cover the aspects needed to be kept in mind while preparing clash free and efficient class schedules for every faculty member. The goal is to generate an optimized solution which will fulfill those constraints while maintaining time efficiency and also reduce the workload of handling this task manually. The proposed algorithm was compared with some base level optimization algorithms to show the better efficiency in terms of accuracy and time. ",
        "url": "https://arxiv.org/abs/2212.02230",
        "authors": [
          "Dibyo Fabian Dofadar",
          "Riyo Hayat Khan",
          "Shafqat Hasan",
          "Towshik Anam Taj",
          "Arif Shakil",
          "Mahbub Majumdar"
        ],
        "subjectives": [
          "Neural and Evolutionary Computing (cs.NE)",
          "Artificial Intelligence (cs.AI)"
        ]
      },
      {
        "id": "arXiv:2212.02233",
        "title": "Wearable-based Human Activity Recognition with Spatio-Temporal Spiking  Neural Networks",
        "abstract": "We study the Human Activity Recognition (HAR) task, which predicts user daily activity based on time series data from wearable sensors. Recently, researchers use end-to-end Artificial Neural Networks (ANNs) to extract the features and perform classification in HAR. However, ANNs pose a huge computation burden on wearable devices and lack temporal feature extraction. In this work, we leverage Spiking Neural Networks (SNNs)--an architecture inspired by biological neurons--to HAR tasks. SNNs allow spatio-temporal extraction of features and enjoy low-power computation with binary spikes. We conduct extensive experiments on three HAR datasets with SNNs, demonstrating that SNNs are on par with ANNs in terms of accuracy while reducing up to 94% energy consumption. The code is publicly available in https://github.com/Intelligent-Computing-Lab-Yale/SNN_HAR ",
        "url": "https://arxiv.org/abs/2212.02233",
        "authors": [
          "Yuhang Li",
          "Ruokai Yin",
          "Hyoungseob Park",
          "Youngeun Kim",
          "Priyadarshini Panda"
        ],
        "subjectives": [
          "Neural and Evolutionary Computing (cs.NE)",
          "Machine Learning (cs.LG)",
          "Signal Processing (eess.SP)"
        ]
      },
      {
        "id": "arXiv:2212.02234",
        "title": "Review of medical data analysis based on spiking neural networks",
        "abstract": "Medical data mainly includes various biomedical signals and medical images, and doctors can make judgments on the physical condition of patients through medical data. However, the interpretation of medical data requires a lot of labor costs and may be misjudged, so many scholars use neural networks and deep learning to classify and study medical data, thereby improving doctors' work efficiency and accuracy, achieving early detection of diseases and early diagnosis, so it has a wide range of application prospects. However, traditional neural networks have disadvantages such as high energy consumption and high latency (slow calculation speed). This paper introduces the research on signal classification and disease diagnosis based on the third-generation neural network - pulse neural network in recent years, using medical data, such as electroencephalogram (EEG), electrocardiogram (ECG), electromyography (EMG), magnetic resonance imaging (MRI), etc., summarizes the advantages and disadvantages of pulse neural networks compared with traditional networks, and looks forward to the future development direction. ",
        "url": "https://arxiv.org/abs/2212.02234",
        "authors": [
          "X. Li",
          "L. Wang",
          "D. Zhao"
        ],
        "subjectives": [
          "Neural and Evolutionary Computing (cs.NE)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.02244",
        "title": "Research on Early Warning and NB-IoT Real-time Monitoring System for  Radiation Source Shedding of Gamma Flaw Detection Machine",
        "abstract": "The system takes the embedded system single chip as the core, and organizes the gamma ray induction module, keying switch, radiation source braid locking mechanism, on-site alarm equipment, NB-IoT communication module, GPS positioning system and other related equipment to realize real-time operators warning and remote alarming to the monitoring platform. Thus, timely management of the fallen radiation source can be realized. What's more, position of radiation source braid is monitored by radiation source braid locking mechanism, gamma-ray induction module and keying switch. This idea solves the bottleneck problem of difficulty in transmitting real-time remote monitoring signal by NB-IoT technology. Single-chip microcontroller is innovatively embedded around the radiation source to monitor whether the source falls off. As a result, when part or all of the source braid are not returned to the storage location of the flaw detector, the gamma ray induction module, keying switch, and radiation source braid locking mechanism remind the operator of the dangerous situation. At the same time, NB-IoT transmission system alarms companies and regulators of safety risks. In conclusion, radiation source can be found timely when it falls off, to avoid radiation accident. Two bottlenecks related to installing GPS positioning system can be solved, that the GPS positioning system is installed on the flaw detector, and the radioactive source is still unable to be controlled, and that the GPS positioning system requires power to transmit the signal back to the regulatory platform or to the platform of the flaw detection enterprise. This article's ideas can solve this problem through the advantages of NB-IoT. ",
        "url": "https://arxiv.org/abs/2212.02244",
        "authors": [
          "Zheng-yang Zhang",
          "Zhi-hui Liu",
          "Rui Zhang",
          "Rong-hua He",
          "Zhe Wang"
        ],
        "subjectives": [
          "Systems and Control (eess.SY)"
        ]
      },
      {
        "id": "arXiv:2212.02263",
        "title": "Counteracting Eavesdropper Attacks Through Reconfigurable Intelligent  Surfaces: A New Threat Model and Secrecy Rate Optimization",
        "abstract": "The potential of Reconfigurable Intelligent Surfaces (RISs) for energy-efficient and performance-boosted wireless communications is recently gaining remarkable research attention, motivating their consideration for various $5$-th Generation (5G) Advanced and beyond applications. In this paper, we consider a Multiple-Input Multiple-Output (MIMO) Physical Layer Security (PLS) system with multiple data streams including one legitimate passive RIS and one malicious passive RIS, with the former being transparent to the multi-antenna eavesdropper and the latter's presence being unknown at the legitimate multi-antenna transceivers. We first present a novel threat model for the RIS-boosted eavesdropping system and design a joint optimization framework for the eavesdropper's receive combining matrix and the reflection coefficients of the malicious RIS. Focusing next on the secrecy rate maximization problem, we present an RIS-empowered PLS scheme that jointly designs the legitimate precoding matrix and number of data streams, the Artificial Noise (AN) covariance matrix, the receive combining matrix, and the reflection coefficients of the legitimate RIS. The proposed optimization algorithms, whose convergence to at least local optimum points is proved, are based on alternating maximization, minorization-maximization, and manifold optimization, including semi-closed form expressions for the optimization variables. Our extensive simulation results for two representative system setups reveal that, in the absence of a legitimate RIS, transceiver spatial filtering and AN are incapable of offering non-zero secrecy rates, even for malicious RISs with small numbers of elements. However, when an $L$-element legitimate RIS is deployed, confidential communication can be safeguarded against eavesdropping systems possessing even more than a $5L$-element malicious RIS. ",
        "url": "https://arxiv.org/abs/2212.02263",
        "authors": [
          "George C. Alexandropoulos",
          "Konstantinos D. Katsanos",
          "Miaowen Wen",
          "Daniel B. da Costa"
        ],
        "subjectives": [
          "Information Theory (cs.IT)"
        ]
      },
      {
        "id": "arXiv:2212.02269",
        "title": "Federated Neural Topic Models",
        "abstract": "Over the last years, topic modeling has emerged as a powerful technique for organizing and summarizing big collections of documents or searching for particular patterns in them. However, privacy concerns arise when cross-analyzing data from different sources is required. Federated topic modeling solves this issue by allowing multiple parties to jointly train a topic model without sharing their data. While several federated approximations of classical topic models do exist, no research has been carried out on their application for neural topic models. To fill this gap, we propose and analyze a federated implementation based on state-of-the-art neural topic modeling implementations, showing its benefits when there is a diversity of topics across the nodes' documents and the need to build a joint model. Our approach is by construction theoretically and in practice equivalent to a centralized approach but preserves the privacy of the nodes. ",
        "url": "https://arxiv.org/abs/2212.02269",
        "authors": [
          "Lorena Calvo-Bartolom\u00e9",
          "Jer\u00f3nimo Arenas-Garc\u00eda"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Computation and Language (cs.CL)"
        ]
      },
      {
        "id": "arXiv:2212.02277",
        "title": "R2FD2: Fast and Robust Matching of Multimodal Remote Sensing Image via  Repeatable Feature Detector and Rotation-invariant Feature Descriptor",
        "abstract": "Automatically identifying feature correspondences between multimodal images is facing enormous challenges because of the significant differences both in radiation and geometry. To address these problems, we propose a novel feature matching method, named R2FD2, that is robust to radiation and rotation differences.Our R2FD2 is conducted in two critical contributions, consisting of a repeatable feature detector and a rotation-invariant feature descriptor. In the first stage, a repeatable feature detector called the Multi-channel Auto-correlation of the Log-Gabor is presented for feature detection, which combines the multi-channel auto-correlation strategy with the Log-Gabor wavelets to detect interest points with high repeatability and uniform distribution. In the second stage, a rotation-invariant feature descriptor is constructed, named the Rotation-invariant Maximum index map of the Log-Gabor, which consists of two components: fast assignment of dominant orientation and construction of feature representation. In the process of fast assignment of dominant orientation, a Rotation-invariant Maximum Index Map is built to address rotation deformations. Then, the proposed RMLG incorporates the rotation-invariant RMIM with the spatial configuration of DAISY to depict a more discriminative feature representation, which improves RMLGs resistance to radiation and rotation variances. ",
        "url": "https://arxiv.org/abs/2212.02277",
        "authors": [
          "Bai Zhu",
          "Chao Yang",
          "Jinkun Dai",
          "Jianwei Fan",
          "Yuanxin Ye"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.02280",
        "title": "GARF:Geometry-Aware Generalized Neural Radiance Field",
        "abstract": "Neural Radiance Field (NeRF) has revolutionized free viewpoint rendering tasks and achieved impressive results. However, the efficiency and accuracy problems hinder its wide applications. To address these issues, we propose Geometry-Aware Generalized Neural Radiance Field (GARF) with a geometry-aware dynamic sampling (GADS) strategy to perform real-time novel view rendering and unsupervised depth estimation on unseen scenes without per-scene optimization. Distinct from most existing generalized NeRFs, our framework infers the unseen scenes on both pixel-scale and geometry-scale with only a few input images. More specifically, our method learns common attributes of novel-view synthesis by an encoder-decoder structure and a point-level learnable multi-view feature fusion module which helps avoid occlusion. To preserve scene characteristics in the generalized model, we introduce an unsupervised depth estimation module to derive the coarse geometry, narrow down the ray sampling interval to proximity space of the estimated surface and sample in expectation maximum position, constituting Geometry-Aware Dynamic Sampling strategy (GADS). Moreover, we introduce a Multi-level Semantic Consistency loss (MSC) to assist more informative representation learning. Extensive experiments on indoor and outdoor datasets show that comparing with state-of-the-art generalized NeRF methods, GARF reduces samples by more than 25\\%, while improving rendering quality and 3D geometry estimation. ",
        "url": "https://arxiv.org/abs/2212.02280",
        "authors": [
          "Yue Shi",
          "Dingyi Rong",
          "Bingbing Ni",
          "Chang Chen",
          "Wenjun Zhang"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.02295",
        "title": "Block Selection Method for Using Feature Norm in Out-of-distribution  Detection",
        "abstract": "Detecting out-of-distribution (OOD) inputs during the inference stage is crucial for deploying neural networks in the real world. Previous methods commonly relied on the output of a network derived from the highly activated feature map. In this study, we first revealed that a norm of the feature map obtained from the other block than the last block can be a better indicator of OOD detection. Motivated by this, we propose a simple framework consisting of FeatureNorm: a norm of the feature map and NormRatio: a ratio of FeatureNorm for ID and OOD to measure the OOD detection performance of each block. In particular, to select the block that provides the largest difference between FeatureNorm of ID and FeatureNorm of OOD, we create Jigsaw puzzle images as pseudo OOD from ID training samples and calculate NormRatio, and the block with the largest value is selected. After the suitable block is selected, OOD detection with the FeatureNorm outperforms other OOD detection methods by reducing FPR95 by up to 52.77% on CIFAR10 benchmark and by up to 48.53% on ImageNet benchmark. We demonstrate that our framework can generalize to various architectures and the importance of block selection, which can improve previous OOD detection methods as well. ",
        "url": "https://arxiv.org/abs/2212.02295",
        "authors": [
          "Yeonguk Yu",
          "Sungho Shin",
          "Seongju Lee",
          "Changhyun Jun",
          "Kyoobin Lee"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.02303",
        "title": "Lossy Compression for Robust Unsupervised Time-Series Anomaly Detection",
        "abstract": "A new Lossy Causal Temporal Convolutional Neural Network Autoencoder for anomaly detection is proposed in this work. Our framework uses a rate-distortion loss and an entropy bottleneck to learn a compressed latent representation for the task. The main idea of using a rate-distortion loss is to introduce representation flexibility that ignores or becomes robust to unlikely events with distinctive patterns, such as anomalies. These anomalies manifest as unique distortion features that can be accurately detected in testing conditions. This new architecture allows us to train a fully unsupervised model that has high accuracy in detecting anomalies from a distortion score despite being trained with some portion of unlabelled anomalous data. This setting is in stark contrast to many of the state-of-the-art unsupervised methodologies that require the model to be only trained on \"normal data\". We argue that this partially violates the concept of unsupervised training for anomaly detection as the model uses an informed decision that selects what is normal from abnormal for training. Additionally, there is evidence to suggest it also effects the models ability at generalisation. We demonstrate that models that succeed in the paradigm where they are only trained on normal data fail to be robust when anomalous data is injected into the training. In contrast, our compression-based approach converges to a robust representation that tolerates some anomalous distortion. The robust representation achieved by a model using a rate-distortion loss can be used in a more realistic unsupervised anomaly detection scheme. ",
        "url": "https://arxiv.org/abs/2212.02303",
        "authors": [
          "Christopher P. Ley",
          "Jorge F. Silva"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Information Theory (cs.IT)"
        ]
      },
      {
        "id": "arXiv:2212.02319",
        "title": "Robust and Accurate Cylinder Triangulation",
        "abstract": "In this paper we present methods for triangulation of infinite cylinders from image line silhouettes. We show numerically that linear estimation of a general quadric surface is inherently a badly posed problem. Instead we propose to constrain the conic section to a circle, and give algebraic constraints on the dual conic, that models this manifold. Using these constraints we derive a fast minimal solver based on three image silhouette lines, that can be used to bootstrap robust estimation schemes such as RANSAC. We also present a constrained least squares solver that can incorporate all available image lines for accurate estimation. The algorithms are tested on both synthetic and real data, where they are shown to give accurate results, compared to previous methods. ",
        "url": "https://arxiv.org/abs/2212.02319",
        "authors": [
          "Anna Gummeson",
          "Magnus Oskarsson"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.02323",
        "title": "Improved Convergence Guarantees for Shallow Neural Networks",
        "abstract": "We continue a long line of research aimed at proving convergence of depth 2 neural networks, trained via gradient descent, to a global minimum. Like in many previous works, our model has the following features: regression with quadratic loss function, fully connected feedforward architecture, RelU activations, Gaussian data instances and network initialization, adversarial labels. It is more general in the sense that we allow both layers to be trained simultaneously and at {\\em different} rates. Our results improve on state-of-the-art [Oymak Soltanolkotabi 20] (training the first layer only) and [Nguyen 21, Section 3.2] (training both layers with Le Cun's initialization). We also report several simple experiments with synthetic data. They strongly suggest that, at least in our model, the convergence phenomenon extends well beyond the ``NTK regime''. ",
        "url": "https://arxiv.org/abs/2212.02323",
        "authors": [
          "Alexander Razborov"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.02340",
        "title": "CBNet: A Plug-and-Play Network for Segmentation-based Scene Text  Detection",
        "abstract": "Recently, segmentation-based methods are quite popular in scene text detection, which mainly contain two steps: text kernel segmentation and expansion. However, the segmentation process only considers each pixel independently, and the expansion process is difficult to achieve a favorable accuracy-speed trade-off. In this paper, we propose a Context-aware and Boundary-guided Network (CBN) to tackle these problems. In CBN, a basic text detector is firstly used to predict initial segmentation results. Then, we propose a context-aware module to enhance text kernel feature representations, which considers both global and local contexts. Finally, we introduce a boundary-guided module to expand enhanced text kernels adaptively with only the pixels on the contours, which not only obtains accurate text boundaries but also keeps high speed, especially on high-resolution output maps. In particular, with a lightweight backbone, the basic detector equipped with our proposed CBN achieves state-of-the-art results on several popular benchmarks, and our proposed CBN can be plugged into several segmentation-based methods. Code will be available on https://github.com/XiiZhao/cbn.pytorch. ",
        "url": "https://arxiv.org/abs/2212.02340",
        "authors": [
          "Xi Zhao",
          "Wei Feng",
          "Zheng Zhang",
          "Jingjing Lv",
          "Xin Zhu",
          "Zhangang Lin",
          "Jinghe Hu",
          "Jingping Shao"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.02346",
        "title": "Accu-Help: A Machine Learning based Smart Healthcare Framework for  Accurate Detection of Obsessive Compulsive Disorder",
        "abstract": "In recent years the importance of Smart Healthcare cannot be overstated. The current work proposed to expand the state-of-art of smart healthcare in integrating solutions for Obsessive Compulsive Disorder (OCD). Identification of OCD from oxidative stress biomarkers (OSBs) using machine learning is an important development in the study of OCD. However, this process involves the collection of OCD class labels from hospitals, collection of corresponding OSBs from biochemical laboratories, integrated and labeled dataset creation, use of suitable machine learning algorithm for designing OCD prediction model, and making these prediction models available for different biochemical laboratories for OCD prediction for unlabeled OSBs. Further, from time to time, with significant growth in the volume of the dataset with labeled samples, redesigning the prediction model is required for further use. The whole process requires distributed data collection, data integration, coordination between the hospital and biochemical laboratory, dynamic machine learning OCD prediction mode design using a suitable machine learning algorithm, and making the machine learning model available for the biochemical laboratories. Keeping all these things in mind, Accu-Help a fully automated, smart, and accurate OCD detection conceptual model is proposed to help the biochemical laboratories for efficient detection of OCD from OSBs. OSBs are classified into three classes: Healthy Individual (HI), OCD Affected Individual (OAI), and Genetically Affected Individual (GAI). The main component of this proposed framework is the machine learning OCD prediction model design. In this Accu-Help, a neural network-based approach is presented with an OCD prediction accuracy of 86 percent. ",
        "url": "https://arxiv.org/abs/2212.02346",
        "authors": [
          "Kabita Patel",
          "Ajaya Kumar Tripathy",
          "Laxmi Narayan Padhy",
          "Sujita Kumar Kar",
          "Susanta Kumar Padhy",
          "Saraju Prasad Mohanty"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.02374",
        "title": "Understanding the Relationship between Over-smoothing and Over-squashing  in Graph Neural Networks",
        "abstract": "Graph Neural Networks (GNNs) have been successfully applied in many applications in computer sciences. Despite the success of deep learning architectures in other domains, deep GNNs still underperform their shallow counterparts. There are many open questions about deep GNNs, but over-smoothing and over-squashing are perhaps the most intriguing issues. When stacking multiple graph convolutional layers, the over-smoothing and over-squashing problems arise and have been defined as the inability of GNNs to learn deep representations and propagate information from distant nodes, respectively. Even though the widespread definitions of both problems are similar, these phenomena have been studied independently. This work strives to understand the underlying relationship between over-smoothing and over-squashing from a topological perspective. We show that both problems are intrinsically related to the spectral gap of the Laplacian of the graph. Therefore, there is a trade-off between these two problems, i.e., we cannot simultaneously alleviate both over-smoothing and over-squashing. We also propose a Stochastic Jost and Liu curvature Rewiring (SJLR) algorithm based on a bound of the Ollivier's Ricci curvature. SJLR is less expensive than previous curvature-based rewiring methods while retaining fundamental properties. Finally, we perform a thorough comparison of SJLR with previous techniques to alleviate over-smoothing or over-squashing, seeking to gain a better understanding of both problems. ",
        "url": "https://arxiv.org/abs/2212.02374",
        "authors": [
          "Jhony H. Giraldo",
          "Fragkiskos D. Malliaros",
          "Thierry Bouwmans"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Artificial Intelligence (cs.AI)"
        ]
      },
      {
        "id": "arXiv:2212.02383",
        "title": "An Approach for Detecting Dynamic Communities in Social Networks",
        "abstract": "Recent developments in the internet and technology have made major advancements in tools that facilitate the collection of social data, opening up thus new opportunities for analyzing social networks. Social network analysis studies the patterns of social relations and aims at discovering the hidden features embedded in the structure of social networks. One of the most important features in social networks is community structure : densely knit groups of individuals. The dynamic nature of interaction in social networks often challenges the detection of such community structures. The contributions in this thesis fall into two categories.The first category highlights the problem of identifying overlapping communities over time. To carry out such analysis, a framework called OLCPM (Online Label propagation and Clique Percolation Method) is proposed. It is an online algorithm based on clique percolation and label propagation methods. OLCPM has two main features : the first one is its ability to discover overlapping communities, while the second is its effectiveness in handling fine-grained temporal net works. As for as the second category is concerned, it emphasizes on the problem of analyzing communities that are embedded at different temporal scales. For example, in networks of interaction such as e-mails or phone calls, individuals are involved in daily as well as occasional conversations. We propose a first method for analyzing communities at multiple temporal scales. Hence, the dynamic network (link streams) is studied at different temporal granularities, and coherent communities (called stable communities) over a period of time are detected at each temporal granularity. The two proposed approaches are validated on both synthetic and real-world datasets. ",
        "url": "https://arxiv.org/abs/2212.02383",
        "authors": [
          "Souaad Boudebza"
        ],
        "subjectives": [
          "Social and Information Networks (cs.SI)",
          "Physics and Society (physics.soc-ph)"
        ]
      },
      {
        "id": "arXiv:2212.02397",
        "title": "PowRL: A Reinforcement Learning Framework for Robust Management of Power  Networks",
        "abstract": "Power grids, across the world, play an important societal and economical role by providing uninterrupted, reliable and transient-free power to several industries, businesses and household consumers. With the advent of renewable power resources and EVs resulting into uncertain generation and highly dynamic load demands, it has become ever so important to ensure robust operation of power networks through suitable management of transient stability issues and localize the events of blackouts. In the light of ever increasing stress on the modern grid infrastructure and the grid operators, this paper presents a reinforcement learning (RL) framework, PowRL, to mitigate the effects of unexpected network events, as well as reliably maintain electricity everywhere on the network at all times. The PowRL leverages a novel heuristic for overload management, along with the RL-guided decision making on optimal topology selection to ensure that the grid is operated safely and reliably (with no overloads). PowRL is benchmarked on a variety of competition datasets hosted by the L2RPN (Learning to Run a Power Network). Even with its reduced action space, PowRL tops the leaderboard in the L2RPN NeurIPS 2020 challenge (Robustness track) at an aggregate level, while also being the top performing agent in the L2RPN WCCI 2020 challenge. Moreover, detailed analysis depicts state-of-the-art performances by the PowRL agent in some of the test scenarios. ",
        "url": "https://arxiv.org/abs/2212.02397",
        "authors": [
          "Anandsingh Chauhan",
          "Mayank Baranwal",
          "Ansuma Basumatary"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Artificial Intelligence (cs.AI)",
          "Systems and Control (eess.SY)",
          "Machine Learning (stat.ML)"
        ]
      },
      {
        "id": "arXiv:2212.02400",
        "title": "Location-Aware Self-Supervised Transformers",
        "abstract": "Pixel-level labels are particularly expensive to acquire. Hence, pretraining is a critical step to improve models on a task like semantic segmentation. However, prominent algorithms for pretraining neural networks use image-level objectives, e.g. image classification, image-text alignment a la CLIP, or self-supervised contrastive learning. These objectives do not model spatial information, which might be suboptimal when finetuning on downstream tasks with spatial reasoning. In this work, we propose to pretrain networks for semantic segmentation by predicting the relative location of image parts. We formulate this task as a classification problem where each patch in a query view has to predict its position relatively to another reference view. We control the difficulty of the task by masking a subset of the reference patch features visible to those of the query. Our experiments show that this location-aware (LOCA) self-supervised pretraining leads to representations that transfer competitively to several challenging semantic segmentation benchmarks. ",
        "url": "https://arxiv.org/abs/2212.02400",
        "authors": [
          "Mathilde Caron",
          "Neil Houlsby",
          "Cordelia Schmid"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.02409",
        "title": "Decoding natural image stimuli from fMRI data with a surface-based  convolutional network",
        "abstract": "Due to the low signal-to-noise ratio and limited resolution of functional MRI data, and the high complexity of natural images, reconstructing a visual stimulus from human brain fMRI measurements is a challenging task. In this work, we propose a novel approach for this task, which we call Cortex2Image, to decode visual stimuli with high semantic fidelity and rich fine-grained detail. In particular, we train a surface-based convolutional network model that maps from brain response to semantic image features first (Cortex2Semantic). We then combine this model with a high-quality image generator (Instance-Conditioned GAN) to train another mapping from brain response to fine-grained image features using a variational approach (Cortex2Detail). Image reconstructions obtained by our proposed method achieve state-of-the-art semantic fidelity, while yielding good fine-grained similarity with the ground-truth stimulus. Our code is available at: https://github.com/zijin-gu/meshconv-decoding.git. ",
        "url": "https://arxiv.org/abs/2212.02409",
        "authors": [
          "Zijin Gu",
          "Keith Jamison",
          "Amy Kuceyeski",
          "Mert Sabuncu"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Machine Learning (cs.LG)",
          "Quantitative Methods (q-bio.QM)"
        ]
      },
      {
        "id": "arXiv:2212.02419",
        "title": "Exploring Graph-aware Multi-View Fusion for Rumor Detection on Social  Media",
        "abstract": "Automatic detecting rumors on social media has become a challenging task. Previous studies focus on learning indicative clues from conversation threads for identifying rumorous information. However, these methods only model rumorous conversation threads from various views but fail to fuse multi-view features very well. In this paper, we propose a novel multi-view fusion framework for rumor representation learning and classification. It encodes the multiple views based on Graph Convolutional Networks (GCN), and leverages Convolutional Neural Networks (CNN) to capture the consistent and complementary information among all views and fuse them together. Experimental results on two public datasets demonstrate that our method outperforms state-of-the-art approaches. ",
        "url": "https://arxiv.org/abs/2212.02419",
        "authors": [
          "Yang Wu",
          "Jing Yang",
          "Xiaojun Zhou",
          "Liming Wang",
          "Zhen Xu"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)"
        ]
      },
      {
        "id": "arXiv:2212.02454",
        "title": "Extending Expressive Access Policies with Privacy Features",
        "abstract": "Authentication, authorization, and trust verification are central parts of an access control system. The conditions for granting access in such a system are collected in access policies. Since access conditions are often complex, dedicated languages -- policy languages -- for defining policies are in use. However, current policy languages are unable to express such conditions having privacy of users in mind. With privacy-preserving technologies, users are enabled to prove information to the access system without revealing it. In this work, we present a generic design for supporting privacy-preserving technologies in policy languages. Our design prevents unnecessary disclosure of sensitive information while still allowing the formulation of expressive rules for access control. For that we make use of zero-knowledge proofs (NIZKs). We demonstrate our design by applying it to the TPL policy language, while using SNARKs. Also, we evaluate the resulting ZK-TPL language and its associated toolchain. Our evaluation shows that for regular-sized credentials communication and verification overhead is negligible. ",
        "url": "https://arxiv.org/abs/2212.02454",
        "authors": [
          "Stefan More",
          "Sebastian Ramacher",
          "Lukas Alber",
          "Marco Herzl"
        ],
        "subjectives": [
          "Cryptography and Security (cs.CR)"
        ]
      },
      {
        "id": "arXiv:2212.02468",
        "title": "Quantized Wasserstein Procrustes Alignment of Word Embedding Spaces",
        "abstract": "Optimal Transport (OT) provides a useful geometric framework to estimate the permutation matrix under unsupervised cross-lingual word embedding (CLWE) models that pose the alignment task as a Wasserstein-Procrustes problem. However, linear programming algorithms and approximate OT solvers via Sinkhorn for computing the permutation matrix come with a significant computational burden since they scale cubically and quadratically, respectively, in the input size. This makes it slow and infeasible to compute OT distances exactly for a larger input size, resulting in a poor approximation quality of the permutation matrix and subsequently a less robust learned transfer function or mapper. This paper proposes an unsupervised projection-based CLWE model called quantized Wasserstein Procrustes (qWP). qWP relies on a quantization step of both the source and target monolingual embedding space to estimate the permutation matrix given a cheap sampling procedure. This approach substantially improves the approximation quality of empirical OT solvers given fixed computational cost. We demonstrate that qWP achieves state-of-the-art results on the Bilingual lexicon Induction (BLI) task. ",
        "url": "https://arxiv.org/abs/2212.02468",
        "authors": [
          "Prince O Aboagye",
          "Yan Zheng",
          "Michael Yeh",
          "Junpeng Wang",
          "Zhongfang Zhuang",
          "Huiyuan Chen",
          "Liang Wang",
          "Wei Zhang",
          "Jeff Phillips"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)"
        ]
      },
      {
        "id": "arXiv:2212.02483",
        "title": "TIDE: Time Derivative Diffusion for Deep Learning on Graphs",
        "abstract": "A prominent paradigm for graph neural networks is based on the message passing framework. In this framework, information communication is realized only between neighboring nodes. The challenge of approaches that use this paradigm is to ensure efficient and accurate \\textit{long distance communication} between nodes, as deep convolutional networks are prone to over-smoothing. In this paper, we present a novel method based on time derivative graph diffusion (TIDE), with a learnable time parameter. Our approach allows to adapt the spatial extent of diffusion across different tasks and network channels, thus enabling medium and long-distance communication efficiently. Furthermore, we show that our architecture directly enables local message passing and thus inherits from the expressive power of local message passing approaches. We show that on widely used graph benchmarks we achieve comparable performance and on a synthetic mesh dataset we outperform state-of-the-art methods like GCN or GRAND by a significant margin. ",
        "url": "https://arxiv.org/abs/2212.02483",
        "authors": [
          "Maximilian Krahn",
          "Maysam Behmanesh",
          "Maks Ovsjanikov"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Social and Information Networks (cs.SI)"
        ]
      },
      {
        "id": "arXiv:2212.02493",
        "title": "Canonical Fields: Self-Supervised Learning of Pose-Canonicalized Neural  Fields",
        "abstract": "Coordinate-based implicit neural networks, or neural fields, have emerged as useful representations of shape and appearance in 3D computer vision. Despite advances however, it remains challenging to build neural fields for categories of objects without datasets like ShapeNet that provide canonicalized object instances that are consistently aligned for their 3D position and orientation (pose). We present Canonical Field Network (CaFi-Net), a self-supervised method to canonicalize the 3D pose of instances from an object category represented as neural fields, specifically neural radiance fields (NeRFs). CaFi-Net directly learns from continuous and noisy radiance fields using a Siamese network architecture that is designed to extract equivariant field features for category-level canonicalization. During inference, our method takes pre-trained neural radiance fields of novel object instances at arbitrary 3D pose, and estimates a canonical field with consistent 3D pose across the entire category. Extensive experiments on a new dataset of 1300 NeRF models across 13 object categories show that our method matches or exceeds the performance of 3D point cloud-based methods. ",
        "url": "https://arxiv.org/abs/2212.02493",
        "authors": [
          "Rohith Agaram",
          "Shaurya Dewan",
          "Rahul Sajnani",
          "Adrien Poulenard",
          "Madhava Krishna",
          "Srinath Sridhar"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.02501",
        "title": "SceneRF: Self-Supervised Monocular 3D Scene Reconstruction with Radiance  Fields",
        "abstract": "In the literature, 3D reconstruction from 2D image has been extensively addressed but often still requires geometrical supervision. In this paper, we propose SceneRF, a self-supervised monocular scene reconstruction method with neural radiance fields (NeRF) learned from multiple image sequences with pose. To improve geometry prediction, we introduce new geometry constraints and a novel probabilistic sampling strategy that efficiently update radiance fields. As the latter are conditioned on a single frame, scene reconstruction is achieved from the fusion of multiple synthesized novel depth views. This is enabled by our spherical-decoder, which allows hallucination beyond the input frame field of view. Thorough experiments demonstrate that we outperform all baselines on all metrics for novel depth views synthesis and scene reconstruction. Our code is available at https://astra-vision.github.io/SceneRF. ",
        "url": "https://arxiv.org/abs/2212.02501",
        "authors": [
          "Anh-Quan Cao",
          "Raoul de Charette"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Artificial Intelligence (cs.AI)",
          "Graphics (cs.GR)",
          "Robotics (cs.RO)"
        ]
      },
      {
        "id": "arXiv:2212.01384",
        "title": "Predicting Drug Repurposing Candidates and Their Mechanisms from A  Biomedical Knowledge Graph",
        "abstract": "Computational drug repurposing is a cost- and time-efficient method to identify new indications of approved or experimental drugs/compounds. It is especially critical for emerging and/or orphan diseases due to its cheaper investment and shorter research cycle compared with traditional wet-lab drug discovery approaches. However, the underlying mechanisms of action between repurposed drugs and their target diseases remain largely unknown, which is still an unsolved issue in existing repurposing methods. As such, computational drug repurposing has not been widely adopted in clinical settings. In this work, based on a massive biomedical knowledge graph, we propose a computational drug repurposing framework that not only predicts the treatment probabilities between drugs and diseases but also predicts the path-based, testable mechanisms of action (MOAs) as their biomedical explanations. Specifically, we utilize the GraphSAGE model in an unsupervised manner to integrate each entity's neighborhood information and employ a Random Forest model to predict the treatment probabilities between pairs of drugs and diseases. Moreover, we train an adversarial actor-critic reinforcement learning model to predict the potential MOA for explaining drug purposing. To encourage the model to find biologically reasonable paths, we utilize the curated molecular interactions of drugs and a PubMed-publication-based concept distance to extract potential drug MOA paths from the knowledge graph as \"demonstration paths\" to guide the model during the process of path-finding. Comprehensive experiments and case studies show that the proposed framework outperforms state-of-the-art baselines in both predictive performance of drug repurposing and explanatory performance of recapitulating human-curated DrugMechDB-based paths. ",
        "url": "https://arxiv.org/abs/2212.01384",
        "authors": [
          "Chunyu Ma",
          "Zhihan Zhou",
          "Han Liu",
          "David Koslicki"
        ],
        "subjectives": [
          "Quantitative Methods (q-bio.QM)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.01446",
        "title": "Downscaling Extreme Rainfall Using Physical-Statistical Generative  Adversarial Learning",
        "abstract": "Modeling the risk of extreme weather events in a changing climate is essential for developing effective adaptation and mitigation strategies. Although the available low-resolution climate models capture different scenarios, accurate risk assessment for mitigation and adaption often demands detail that they typically cannot resolve. Here, we develop a dynamic data-driven downscaling (super-resolution) method that incorporates physics and statistics in a generative framework to learn the fine-scale spatial details of rainfall. Our method transforms coarse-resolution ($0.25^{\\circ} \\times 0.25^{\\circ}$) climate model outputs into high-resolution ($0.01^{\\circ} \\times 0.01^{\\circ}$) rainfall fields while efficaciously quantifying uncertainty. Results indicate that the downscaled rainfall fields closely match observed spatial fields and their risk distributions. ",
        "url": "https://arxiv.org/abs/2212.01446",
        "authors": [
          "Anamitra Saha",
          "Sai Ravela"
        ],
        "subjectives": [
          "Atmospheric and Oceanic Physics (physics.ao-ph)",
          "Machine Learning (cs.LG)",
          "Applications (stat.AP)"
        ]
      },
      {
        "id": "arXiv:2212.01480",
        "title": "An efficient numerical algorithm for the moment neural activation",
        "abstract": "Derived from spiking neuron models via the diffusion approximation, the moment activation (MA) faithfully captures the nonlinear coupling of correlated neural variability. However, numerical evaluation of the MA faces significant challenges due to a number of ill-conditioned Dawson-like functions. By deriving asymptotic expansions of these functions, we develop an efficient numerical algorithm for evaluating the MA and its derivatives ensuring reliability, speed, and accuracy. We also provide exact analytical expressions for the MA in the weak fluctuation limit. Powered by this efficient algorithm, the MA may serve as an effective tool for investigating the dynamics of correlated neural variability in large-scale spiking neural circuits. ",
        "url": "https://arxiv.org/abs/2212.01480",
        "authors": [
          "Yang Qi"
        ],
        "subjectives": [
          "Biological Physics (physics.bio-ph)",
          "Numerical Analysis (math.NA)"
        ]
      },
      {
        "id": "arXiv:2212.01518",
        "title": "Hedging against Complexity: Distributionally Robust Optimization with  Parametric Approximation",
        "abstract": "Empirical risk minimization (ERM) and distributionally robust optimization (DRO) are popular approaches for solving stochastic optimization problems that appear in operations management and machine learning. Existing generalization error bounds for these methods depend on either the complexity of the cost function or dimension of the uncertain parameters; consequently, the performance of these methods is poor for high-dimensional problems with objective functions under high complexity. We propose a simple approach in which the distribution of uncertain parameters is approximated using a parametric family of distributions. This mitigates both sources of complexity; however, it introduces a model misspecification error. We show that this new source of error can be controlled by suitable DRO formulations. Our proposed parametric DRO approach has significantly improved generalization bounds over existing ERM / DRO methods and parametric ERM for a wide variety of settings. Our method is particularly effective under distribution shifts. We also illustrate the superior performance of our approach on both synthetic and real-data portfolio optimization and regression tasks. ",
        "url": "https://arxiv.org/abs/2212.01518",
        "authors": [
          "Garud Iyengar",
          "Henry Lam",
          "Tianyu Wang"
        ],
        "subjectives": [
          "Optimization and Control (math.OC)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.01554",
        "title": "Distributionally Robust Lyapunov Function Search Under Uncertainty",
        "abstract": "This paper develops methods for proving Lyapunov stability of dynamical systems subject to disturbances with an unknown distribution. We assume only a finite set of disturbance samples is available and that the true online disturbance realization may be drawn from a different distribution than the given samples. We formulate an optimization problem to search for a sum-of-squares (SOS) Lyapunov function and introduce a distributionally robust version of the Lyapunov function derivative constraint. We show that this constraint may be reformulated as several SOS constraints, ensuring that the search for a Lyapunov function remains in the class of SOS polynomial optimization problems. For general systems, we provide a distributionally robust chance-constrained formulation for neural network Lyapunov function search. Simulations demonstrate the validity and efficiency of either formulation on non-linear uncertain dynamical systems. ",
        "url": "https://arxiv.org/abs/2212.01554",
        "authors": [
          "Kehan Long",
          "Yinzhuang Yi",
          "Jorge Cortes",
          "Nikolay Atanasov"
        ],
        "subjectives": [
          "Optimization and Control (math.OC)",
          "Robotics (cs.RO)"
        ]
      },
      {
        "id": "arXiv:2212.01577",
        "title": "A Domain-specific Perceptual Metric via Contrastive Self-supervised  Representation: Applications on Natural and Medical Images",
        "abstract": "Quantifying the perceptual similarity of two images is a long-standing problem in low-level computer vision. The natural image domain commonly relies on supervised learning, e.g., a pre-trained VGG, to obtain a latent representation. However, due to domain shift, pre-trained models from the natural image domain might not apply to other image domains, such as medical imaging. Notably, in medical imaging, evaluating the perceptual similarity is exclusively performed by specialists trained extensively in diverse medical fields. Thus, medical imaging remains devoid of task-specific, objective perceptual measures. This work answers the question: Is it necessary to rely on supervised learning to obtain an effective representation that could measure perceptual similarity, or is self-supervision sufficient? To understand whether recent contrastive self-supervised representation (CSR) may come to the rescue, we start with natural images and systematically evaluate CSR as a metric across numerous contemporary architectures and tasks and compare them with existing methods. We find that in the natural image domain, CSR behaves on par with the supervised one on several perceptual tests as a metric, and in the medical domain, CSR better quantifies perceptual similarity concerning the experts' ratings. We also demonstrate that CSR can significantly improve image quality in two image synthesis tasks. Finally, our extensive results suggest that perceptuality is an emergent property of CSR, which can be adapted to many image domains without requiring annotations. ",
        "url": "https://arxiv.org/abs/2212.01577",
        "authors": [
          "Hongwei Bran Li",
          "Chinmay Prabhakar",
          "Suprosanna Shit",
          "Johannes Paetzold",
          "Tamaz Amiranashvili",
          "Jianguo Zhang",
          "Daniel Rueckert",
          "Juan Eugenio Iglesias",
          "Benedikt Wiestler",
          "Bjoern Menze"
        ],
        "subjectives": [
          "Image and Video Processing (eess.IV)",
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Multimedia (cs.MM)"
        ]
      },
      {
        "id": "arXiv:2212.01625",
        "title": "Power network optimization: a quantum approach",
        "abstract": "Optimization of electricity surplus is a crucial element for transmission power networks to reduce costs and efficiently use the available electricity across the network. In this paper we showed how to optimize such a network with quantum annealing. First, we define the QUBO problem for the partitioning of the network, and test the implementation on purely quantum and hybrid architectures. We then solve the problem on the D-Wave hybrid CQM and BQM solvers, as well as on classical solvers available on Azure Quantum cloud. Finally, we show that the hybrid approaches overperform the classical methods in terms of quality of the solution, as the value of the objective function of the quantum solutions is found to be always lower than with the classical approaches across a set of different problem size. ",
        "url": "https://arxiv.org/abs/2212.01625",
        "authors": [
          "Giuseppe Colucci",
          "Stan van der Linde",
          "Frank Phillipson"
        ],
        "subjectives": [
          "Quantum Physics (quant-ph)",
          "Emerging Technologies (cs.ET)",
          "Optimization and Control (math.OC)"
        ]
      },
      {
        "id": "arXiv:2212.01657",
        "title": "SINR Coverage Enhancement of 6G UAV-Assisted Networks Deploying IRS",
        "abstract": "The objective of the work is to enhance the signal-to-interference-plus-noise ratio (SINR) coverage probability in UAV-assisted 6G wireless networks. Therefore, the work analyzed and compared the coverage probability of conventional and IRS-assisted UAV communications models. The research obtained that the employment of IRS in UAV-assisted wireless networks enhances the SINR coverage probability significantly. Moreover, the installment of IRS in UAV networks reduces the energy consumption of the network. ",
        "url": "https://arxiv.org/abs/2212.01657",
        "authors": [
          "Mobasshir Mahbub",
          "Raed M. Shubair"
        ],
        "subjectives": [
          "Signal Processing (eess.SP)",
          "Information Theory (cs.IT)"
        ]
      },
      {
        "id": "arXiv:2212.01661",
        "title": "Unsupervised Fine-Tuning Data Selection for ASR Using Self-Supervised  Speech Models",
        "abstract": "Self-supervised learning (SSL) has been able to leverage unlabeled data to boost the performance of automatic speech recognition (ASR) models when we have access to only a small amount of transcribed speech data. However, this raises the question of which subset of the available unlabeled data should be selected for transcription. Our work investigates different unsupervised data selection techniques for fine-tuning the HuBERT model under a limited transcription budget. We investigate the impact of speaker diversity, gender bias, and topic diversity on the downstream ASR performance. We also devise two novel techniques for unsupervised data selection: pre-training loss based data selection and the perplexity of byte pair encoded clustered units (PBPE) and we show how these techniques compare to pure random data selection. Finally, we analyze the correlations between the inherent characteristics of the selected fine-tuning subsets as well as how these characteristics correlate with the resultant word error rate. We demonstrate the importance of token diversity, speaker diversity, and topic diversity in achieving the best performance in terms of WER. ",
        "url": "https://arxiv.org/abs/2212.01661",
        "authors": [
          "Reem Gody",
          "David Harwath"
        ],
        "subjectives": [
          "Audio and Speech Processing (eess.AS)",
          "Computation and Language (cs.CL)",
          "Sound (cs.SD)"
        ]
      },
      {
        "id": "arXiv:2212.01694",
        "title": "A Quantum Overlay Network for Efficient Entanglement Distribution",
        "abstract": "Distributing quantum entanglements over long distances is essential for the realization of a global scale quantum Internet. Most of the prior work and proposals assume an on-demand distribution of entanglements which may result in significant network resource under-utilization. In this work, we introduce Quantum Overlay Networks (QONs) for efficient entanglement distribution in quantum networks. When the demand to create end-to-end user entanglements is low, QONs can generate and store maximally entangled Bell pairs (EPR pairs) at specific overlay storage nodes of the network. Later, during peak demands, requests can be served by performing entanglement swaps either over a direct path from the network or over a path using the storage nodes. We solve the link entanglement and storage resource allocation problem in such a QON using a centralized optimization framework. We evaluate the performance of our proposed QON architecture over a wide number of network topologies under various settings using extensive simulation experiments. Our results demonstrate that QONs fare well by a factor of 40% with respect to meeting surge and changing demands compared to traditional non-overlay proposals. QONs also show significant improvement in terms of average entanglement request service delay over non-overlay approaches. ",
        "url": "https://arxiv.org/abs/2212.01694",
        "authors": [
          "Shahrooz Pouryousef",
          "Nitish K. Panigrahy",
          "Don Towsley"
        ],
        "subjectives": [
          "Quantum Physics (quant-ph)",
          "Networking and Internet Architecture (cs.NI)",
          "Performance (cs.PF)"
        ]
      },
      {
        "id": "arXiv:2212.01705",
        "title": "Breaking Down the Lockdown: The Causal Effect of Stay-At-Home Mandates  on Uncertainty and Sentiments during the COVID-19 Pandemic",
        "abstract": "We study the causal effects of lockdown measures on uncertainty and sentiments on Twitter. By exploiting the quasi-experimental setting induced by the first Western COVID-19 lockdown - the unexpected lockdown implemented in Northern Italy in February 2020 - we measure changes in public uncertainty and sentiment expressed on daily pre and post-lockdown tweets geolocalized inside and in the proximity of the lockdown areas. Using natural language processing, including dictionary-based methods and deep learning models, we classify each tweet across four categories - economics, health, politics and lockdown policy - to identify in which areas uncertainty and sentiments concentrate. Using a Difference-in-Difference analysis, we show that areas under lockdown depict lower uncertainty around the economy and the stay-at-home mandate itself. This surprising result likely stems from an informational asymmetry channel, for which treated individuals adjusts their expectations once the policy is in place, while uncertainty persists around the untreated. However, we also find that the lockdown comes at a cost as political sentiments worsen. ",
        "url": "https://arxiv.org/abs/2212.01705",
        "authors": [
          "C. Biliotti",
          "F.J. Bargagli-Stoffi",
          "N. Fraccaroli",
          "M. Puliga",
          "M. Riccaboni"
        ],
        "subjectives": [
          "Applications (stat.AP)",
          "Social and Information Networks (cs.SI)",
          "General Economics (econ.GN)"
        ]
      },
      {
        "id": "arXiv:2212.01761",
        "title": "A PM2.5 concentration prediction framework with vehicle tracking system:  From cause to effect",
        "abstract": "Air pollution is an emerging problem that needs to be solved especially in developed and developing countries. In Vietnam, air pollution is also a concerning issue in big cities such as Hanoi and Ho Chi Minh cities where air pollution comes mostly from vehicles such as cars and motorbikes. In order to tackle the problem, the paper focuses on developing a solution that can estimate the emitted PM2.5 pollutants by counting the number of vehicles in the traffic. We first investigated among the recent object detection models and developed our own traffic surveillance system. The observed traffic density showed a similar trend to the measured PM2.5 with a certain lagging in time, suggesting a relation between traffic density and PM2.5. We further express this relationship with a mathematical model which can estimate the PM2.5 value based on the observed traffic density. The estimated result showed a great correlation with the measured PM2.5 plots in the urban area context. ",
        "url": "https://arxiv.org/abs/2212.01761",
        "authors": [
          "Chuong D. Le",
          "Hoang V. Pham",
          "Duy A. Pham",
          "An D. Le",
          "Hien B. Vo"
        ],
        "subjectives": [
          "Physics and Society (physics.soc-ph)",
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.01816",
        "title": "Joint graph learning from Gaussian observations in the presence of  hidden nodes",
        "abstract": "Graph learning problems are typically approached by focusing on learning the topology of a single graph when signals from all nodes are available. However, many contemporary setups involve multiple related networks and, moreover, it is often the case that only a subset of nodes is observed while the rest remain hidden. Motivated by this, we propose a joint graph learning method that takes into account the presence of hidden (latent) variables. Intuitively, the presence of the hidden nodes renders the inference task ill-posed and challenging to solve, so we overcome this detrimental influence by harnessing the similarity of the estimated graphs. To that end, we assume that the observed signals are drawn from a Gaussian Markov random field with latent variables and we carefully model the graph similarity among hidden (latent) nodes. Then, we exploit the structure resulting from the previous considerations to propose a convex optimization problem that solves the joint graph learning task by providing a regularized maximum likelihood estimator. Finally, we compare the proposed algorithm with different baselines and evaluate its performance over synthetic and real-world graphs. ",
        "url": "https://arxiv.org/abs/2212.01816",
        "authors": [
          "Samuel Rey",
          "Madeline Navarro",
          "Andrei Buciulea",
          "Santiago Segarra",
          "Antonio G. Marques"
        ],
        "subjectives": [
          "Signal Processing (eess.SP)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.01825",
        "title": "MouseGAN++: Unsupervised Disentanglement and Contrastive Representation  for Multiple MRI Modalities Synthesis and Structural Segmentation of Mouse  Brain",
        "abstract": "Segmenting the fine structure of the mouse brain on magnetic resonance (MR) images is critical for delineating morphological regions, analyzing brain function, and understanding their relationships. Compared to a single MRI modality, multimodal MRI data provide complementary tissue features that can be exploited by deep learning models, resulting in better segmentation results. However, multimodal mouse brain MRI data is often lacking, making automatic segmentation of mouse brain fine structure a very challenging task. To address this issue, it is necessary to fuse multimodal MRI data to produce distinguished contrasts in different brain structures. Hence, we propose a novel disentangled and contrastive GAN-based framework, named MouseGAN++, to synthesize multiple MR modalities from single ones in a structure-preserving manner, thus improving the segmentation performance by imputing missing modalities and multi-modality fusion. Our results demonstrate that the translation performance of our method outperforms the state-of-the-art methods. Using the subsequently learned modality-invariant information as well as the modality-translated images, MouseGAN++ can segment fine brain structures with averaged dice coefficients of 90.0% (T2w) and 87.9% (T1w), respectively, achieving around +10% performance improvement compared to the state-of-the-art algorithms. Our results demonstrate that MouseGAN++, as a simultaneous image synthesis and segmentation method, can be used to fuse cross-modality information in an unpaired manner and yield more robust performance in the absence of multimodal data. We release our method as a mouse brain structural segmentation tool for free academic usage at https://github.com/yu02019. ",
        "url": "https://arxiv.org/abs/2212.01825",
        "authors": [
          "Ziqi Yu",
          "Xiaoyang Han",
          "Shengjie Zhang",
          "Jianfeng Feng",
          "Tingying Peng",
          "Xiao-Yong Zhang"
        ],
        "subjectives": [
          "Image and Video Processing (eess.IV)",
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.01953",
        "title": "Context-aware multi-head self-attentional neural network model for next  location prediction",
        "abstract": "Accurate activity location prediction is a crucial component of many mobility applications and is particularly required to develop personalized, sustainable transportation systems. Despite the widespread adoption of deep learning models, next location prediction models lack a comprehensive discussion and integration of mobility-related spatio-temporal contexts. Here, we utilize a multi-head self-attentional (MHSA) neural network that learns location transition patterns from historical location visits, their visit time and activity duration, as well as their surrounding land use functions, to infer an individual's next location. Specifically, we adopt point-of-interest data and latent Dirichlet allocation for representing locations' land use contexts at multiple spatial scales, generate embedding vectors of the spatio-temporal features, and learn to predict the next location with an MHSA network. Through experiments on two large-scale GNSS tracking datasets, we demonstrate that the proposed model outperforms other state-of-the-art prediction models, and reveal the contribution of various spatio-temporal contexts to the model's performance. Moreover, we find that the model trained on population data achieves higher prediction performance with fewer parameters than individual-level models due to learning from collective movement patterns. We also reveal mobility conducted in the recent past and one week before has the largest influence on the current prediction, showing that learning from a subset of the historical mobility is sufficient to obtain an accurate location prediction result. We believe that the proposed model is vital for context-aware mobility prediction. The gained insights will help to understand location prediction models and promote their implementation for mobility applications. ",
        "url": "https://arxiv.org/abs/2212.01953",
        "authors": [
          "Ye Hong",
          "Yatao Zhang",
          "Konrad Schindler",
          "Martin Raubal"
        ],
        "subjectives": [
          "Physics and Society (physics.soc-ph)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.02033",
        "title": "Towards Generating Diverse Audio Captions via Adversarial Training",
        "abstract": "Automated audio captioning is a cross-modal translation task for describing the content of audio clips with natural language sentences. This task has attracted increasing attention and substantial progress has been made in recent years. Captions generated by existing models are generally faithful to the content of audio clips, however, these machine-generated captions are often deterministic (e.g., generating a fixed caption for a given audio clip), simple (e.g., using common words and simple grammar), and generic (e.g., generating the same caption for similar audio clips). When people are asked to describe the content of an audio clip, different people tend to focus on different sound events and describe an audio clip diversely from various aspects using distinct words and grammar. We believe that an audio captioning system should have the ability to generate diverse captions, either for a fixed audio clip, or across similar audio clips. To this end, we propose an adversarial training framework based on a conditional generative adversarial network (C-GAN) to improve diversity of audio captioning systems. A caption generator and two hybrid discriminators compete and are learned jointly, where the caption generator can be any standard encoder-decoder captioning model used to generate captions, and the hybrid discriminators assess the generated captions from different criteria, such as their naturalness and semantics. We conduct experiments on the Clotho dataset. The results show that our proposed model can generate captions with better diversity as compared to state-of-the-art methods. ",
        "url": "https://arxiv.org/abs/2212.02033",
        "authors": [
          "Xinhao Mei",
          "Xubo Liu",
          "Jianyuan Sun",
          "Mark D. Plumbley",
          "Wenwu Wang"
        ],
        "subjectives": [
          "Audio and Speech Processing (eess.AS)",
          "Artificial Intelligence (cs.AI)",
          "Multimedia (cs.MM)",
          "Sound (cs.SD)"
        ]
      },
      {
        "id": "arXiv:2212.02099",
        "title": "LMEC: Learnable Multiplicative Absolute Position Embedding Based  Conformer for Speech Recognition",
        "abstract": "This paper proposes a Learnable Multiplicative absolute position Embedding based Conformer (LMEC). It contains a kernelized linear attention (LA) module called LMLA to solve the time-consuming problem for long sequence speech recognition as well as an alternative to the FFN structure. First, the ELU function is adopted as the kernel function of our proposed LA module. Second, we propose a novel Learnable Multiplicative Absolute Position Embedding (LM-APE) based re-weighting mechanism that can reduce the well-known quadratic temporal-space complexity of softmax self-attention. Third, we use Gated Linear Units (GLU) to substitute the Feed Forward Network (FFN) for better performance. Extensive experiments have been conducted on the public LibriSpeech datasets. Compared to the Conformer model with cosFormer style linear attention, our proposed method can achieve up to 0.63% word-error-rate improvement on test-other and improve the inference speed by up to 13% (left product) and 33% (right product) on the LA module. ",
        "url": "https://arxiv.org/abs/2212.02099",
        "authors": [
          "Yuguang Yang",
          "Yu Pan",
          "Jingjing Yin",
          "Heng Lu"
        ],
        "subjectives": [
          "Audio and Speech Processing (eess.AS)",
          "Sound (cs.SD)"
        ]
      },
      {
        "id": "arXiv:2212.02105",
        "title": "Matrix factorization with neural networks",
        "abstract": "Matrix factorization is an important mathematical problem encountered in the context of dictionary learning, recommendation systems and machine learning. We introduce a new `decimation' scheme that maps it to neural network models of associative memory and provide a detailed theoretical analysis of its performance, showing that decimation is able to factorize extensive-rank matrices and to denoise them efficiently. We introduce a decimation algorithm based on ground-state search of the neural network, which shows performances that match the theoretical prediction. ",
        "url": "https://arxiv.org/abs/2212.02105",
        "authors": [
          "Francesco Camilli",
          "Marc M\u00e9zard"
        ],
        "subjectives": [
          "Disordered Systems and Neural Networks (cond-mat.dis-nn)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.02201",
        "title": "Relative Timing Information and Orthology in Evolutionary Scenarios",
        "abstract": "Evolutionary scenarios describing the evolution of a family of genes within a collection of species comprise the mapping of the vertices of a gene tree $T$ to vertices and edges of a species tree $S$. The relative timing of the last common ancestors of two extant genes (leaves of $T$) and the last common ancestors of the two species (leaves of $S$) in which they reside is indicative of horizontal gene transfers (HGT) and ancient duplications. Orthologous gene pairs, on the other hand, require that their last common ancestors coincides with a corresponding speciation event. The relative timing information of gene and species divergences is captured by three colored graphs that have the extant genes as vertices and the species in which the genes are found as vertex colors: the equal-divergence-time (EDT) graph, the later-divergence-time (LDT) graph and the prior-divergence-time (PDT) graph, which together form an edge partition of the complete graph. Here we give a complete characterization in terms of informative and forbidden triples that can be read off the three graphs and provide a polynomial time algorithm for constructing an evolutionary scenario that explains the graphs, provided such a scenario exists. While both LDT and PDT graphs are cographs, this is not true for the EDT graph in general. We show that every EDT graph is perfect. While the information about LDT and PDT graphs is necessary to recognize EDT graphs in polynomial-time for general scenarios, this extra information can be dropped in the HGT-free case. However, recognition of EDT graphs without knowledge of putative LDT and PDT graphs is NP-complete for general scenarios. We finally connect the EDT graph to the alternative definitions of orthology that have been proposed for scenarios with horizontal gene transfer. With one exception, the corresponding graphs are shown to be colored cographs. ",
        "url": "https://arxiv.org/abs/2212.02201",
        "authors": [
          "David Schaller",
          "Tom Hartmann",
          "Manuel Lafond",
          "Nicolas Wieseke",
          "Marc Hellmuth",
          "Peter F. Stadler"
        ],
        "subjectives": [
          "Populations and Evolution (q-bio.PE)",
          "Computational Complexity (cs.CC)",
          "Discrete Mathematics (cs.DM)",
          "Combinatorics (math.CO)"
        ]
      },
      {
        "id": "arXiv:2212.02223",
        "title": "Limitations on approximation by deep and shallow neural networks",
        "abstract": "We prove Carl's type inequalities for the error of approximation of compact sets K by deep and shallow neural networks. This in turn gives lower bounds on how well we can approximate the functions in K when requiring the approximants to come from outputs of such networks. Our results are obtained as a byproduct of the study of the recently introduced Lipschitz widths. ",
        "url": "https://arxiv.org/abs/2212.02223",
        "authors": [
          "Guergana Petrova",
          "Przemys\u0142aw Wojtaszczyk"
        ],
        "subjectives": [
          "Machine Learning (stat.ML)",
          "Machine Learning (cs.LG)",
          "Functional Analysis (math.FA)"
        ]
      },
      {
        "id": "arXiv:2212.02226",
        "title": "Inferring latent neural sources via deep transcoding of simultaneously  acquired EEG and fMRI",
        "abstract": "Simultaneous EEG-fMRI is a multi-modal neuroimaging technique that provides complementary spatial and temporal resolution. Challenging has been developing principled and interpretable approaches for fusing the modalities, specifically approaches enabling inference of latent source spaces representative of neural activity. In this paper, we address this inference problem within the framework of transcoding -- mapping from a specific encoding (modality) to a decoding (the latent source space) and then encoding the latent source space to the other modality. Specifically, we develop a symmetric method consisting of a cyclic convolutional transcoder that transcodes EEG to fMRI and vice versa. Without any prior knowledge of either the hemodynamic response function or lead field matrix, the complete data-driven method exploits the temporal and spatial relationships between the modalities and latent source spaces to learn these mappings. We quantify, for both the simulated and real EEG-fMRI data, how well the modalities can be transcoded from one to another as well as the source spaces that are recovered, all evaluated on unseen data. In addition to enabling a new way to symmetrically infer a latent source space, the method can also be seen as low-cost computational neuroimaging -- i.e. generating an 'expensive' fMRI BOLD image from 'low cost' EEG data. ",
        "url": "https://arxiv.org/abs/2212.02226",
        "authors": [
          "Xueqing Liu",
          "Tao Tu",
          "Paul Sajda"
        ],
        "subjectives": [
          "Neurons and Cognition (q-bio.NC)",
          "Artificial Intelligence (cs.AI)",
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.02229",
        "title": "GPS++: An Optimised Hybrid MPNN/Transformer for Molecular Property  Prediction",
        "abstract": "This technical report presents GPS++, a top-3 finalist method in Open Graph Benchmark Large-Scale Challenge (OGB-LSC 2022) for PCQM4Mv2 molecular property prediction task. Our approach implements several key principles from the prior literature. At its core our GPS++ method is a hybrid MPNN/Transformer model that incorporates 3D atom positions and an auxiliary denoising task. The effectiveness of GPS++ is demonstrated by achieving 0.0719 mean absolute error on the independent test-challenge PCQM4Mv2 split. Thanks to Graphcore IPU acceleration, GPS++ scales to deep architectures (16 layers), training at 3 minutes per epoch, and large ensemble (112 models), completing the final predictions in 1 hour 32 minutes, well under the 4 hour inference budget allocated. Our implementation is publicly available at: https://github.com/graphcore/ogb-lsc-pcqm4mv2. ",
        "url": "https://arxiv.org/abs/2212.02229",
        "authors": [
          "Dominic Masters",
          "Josef Dean",
          "Kerstin Klaser",
          "Zhiyi Li",
          "Sam Maddrell-Mander",
          "Adam Sanders",
          "Hatem Helal",
          "Deniz Beker",
          "Ladislav Ramp\u00e1\u0161ek",
          "Dominique Beaini"
        ],
        "subjectives": [
          "Quantitative Methods (q-bio.QM)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.02251",
        "title": "Multiscale Graph Neural Networks for Protein Residue Contact Map  Prediction",
        "abstract": "Machine learning (ML) is revolutionizing protein structural analysis, including an important subproblem of predicting protein residue contact maps, i.e., which amino-acid residues are in close spatial proximity given the amino-acid sequence of a protein. Despite recent progresses in ML-based protein contact prediction, predicting contacts with a wide range of distances (commonly classified into short-, medium- and long-range contacts) remains a challenge. Here, we propose a multiscale graph neural network (GNN) based approach taking a cue from multiscale physics simulations, in which a standard pipeline involving a recurrent neural network (RNN) is augmented with three GNNs to refine predictive capability for short-, medium- and long-range residue contacts, respectively. Test results on the ProteinNet dataset show improved accuracy for contacts of all ranges using the proposed multiscale RNN+GNN approach over the conventional approach, including the most challenging case of long-range contact prediction. ",
        "url": "https://arxiv.org/abs/2212.02251",
        "authors": [
          "Kuang Liu",
          "Rajiv K. Kalia",
          "Xinlian Liu",
          "Aiichiro Nakano",
          "Ken-ichi Nomura",
          "Priya Vashishta",
          "Rafael Zamora-Resendizc"
        ],
        "subjectives": [
          "Quantitative Methods (q-bio.QM)",
          "Artificial Intelligence (cs.AI)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.02392",
        "title": "Maximum entropy network states for coalescence processes",
        "abstract": "Complex network states are characterized by the interplay between system's structure and dynamics. One way to represent such states is by means of network density matrices, whose von Neumann entropy characterizes the number of distinct microstates compatible with given topology and dynamical evolution. In this Letter, we propose a maximum entropy principle to characterize network states for systems with heterogeneous, generally correlated, connectivity patterns and non-trivial dynamics. We focus on three distinct coalescence processes, widely encountered in the analysis of empirical interconnected systems, and characterize their entropy and transitions between distinct dynamical regimes across distinct temporal scales. Our framework allows one to study the statistical physics of systems that aggregate, such as in transportation infrastructures serving the same geographic area, or correlate, such as inter-brain synchrony arising in organisms that socially interact, and active matter that swarm or synchronize. ",
        "url": "https://arxiv.org/abs/2212.02392",
        "authors": [
          "Arsham Ghavasieh",
          "Manlio De Domenico"
        ],
        "subjectives": [
          "Physics and Society (physics.soc-ph)",
          "Statistical Mechanics (cond-mat.stat-mech)",
          "Information Theory (cs.IT)"
        ]
      },
      {
        "id": "arXiv:2212.02402",
        "title": "A Network Theory Investigation into the Altered Resting State Functional  Connectivity in Attention-Deficit Hyperactivity Disorder",
        "abstract": "In the last two decades, functional magnetic resonance imaging (fMRI) has emerged as one of the most effective technologies in clinical research of the human brain. fMRI allows researchers to study healthy and pathological brains while they perform various neuropsychological functions. Beyond task-related activations, the human brain has some intrinsic activity at a task-negative (resting) state that surprisingly consumes a lot of energy to support communication among neurons. Recent neuroimaging research has also seen an increase in modeling and analyzing brain activity in terms of a graph or network. Since graph models facilitate a systems-theoretic explanation of the brain, they have become increasingly relevant with advances in network science and the popularization of complex systems theory. The purpose of this study is to look into the abnormalities in resting brain functions in adults with Attention Deficit Hyperactivity Disorder (ADHD). The primary goal is to investigate resting-state functional connectivity (FC), which can be construed as a significant temporal coincidence in blood-oxygen-level dependent (BOLD) signals between functionally related brain regions in the absence of any stimulus or task. When compared to healthy controls, ADHD patients have lower average connectivity in the Supramarginal Gyrus and Superior Parietal Lobule, but higher connectivity in the Lateral Occipital Cortex and Inferior Temporal Gyrus. We also hypothesize that the network organization of default mode and dorsal attention regions is abnormal in ADHD patients. ",
        "url": "https://arxiv.org/abs/2212.02402",
        "authors": [
          "Sadi Md. Redwan",
          "Md Palash Uddin",
          "Muhammad Imran Sharif",
          "Anwaar Ulhaq"
        ],
        "subjectives": [
          "Neurons and Cognition (q-bio.NC)",
          "Machine Learning (cs.LG)",
          "Signal Processing (eess.SP)"
        ]
      },
      {
        "id": "arXiv:2212.02435",
        "title": "Observational and Interventional Causal Learning for Regret-Minimizing  Control",
        "abstract": "We explore how observational and interventional causal discovery methods can be combined. A state-of-the-art observational causal discovery algorithm for time series capable of handling latent confounders and contemporaneous effects, called LPCMCI, is extended to profit from casual constraints found through randomized control trials. Numerical results show that, given perfect interventional constraints, the reconstructed structural causal models (SCMs) of the extended LPCMCI allow 84.6% of the time for the optimal prediction of the target variable. The implementation of interventional and observational causal discovery is modular, allowing causal constraints from other sources. The second part of this thesis investigates the question of regret minimizing control by simultaneously learning a causal model and planning actions through the causal model. The idea is that an agent to optimize a measured variable first learns the system's mechanics through observational causal discovery. The agent then intervenes on the most promising variable with randomized values allowing for the exploitation and generation of new interventional data. The agent then uses the interventional data to enhance the causal model further, allowing improved actions the next time. The extended LPCMCI can be favorable compared to the original LPCMCI algorithm. The numerical results show that detecting and using interventional constraints leads to reconstructed SCMs that allow 60.9% of the time for the optimal prediction of the target variable in contrast to the baseline of 53.6% when using the original LPCMCI algorithm. Furthermore, the induced average regret decreases from 1.2 when using the original LPCMCI algorithm to 1.0 when using the extended LPCMCI algorithm with interventional discovery. ",
        "url": "https://arxiv.org/abs/2212.02435",
        "authors": [
          "Christian Reiser"
        ],
        "subjectives": [
          "Machine Learning (stat.ML)",
          "Artificial Intelligence (cs.AI)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2212.02457",
        "title": "Blessings and Curses of Covariate Shifts: Adversarial Learning Dynamics,  Directional Convergence, and Equilibria",
        "abstract": "Covariate distribution shifts and adversarial perturbations present robustness challenges to the conventional statistical learning framework: seemingly small unconceivable shifts in the test covariate distribution can significantly affect the performance of the statistical model learned based on the training distribution. The model performance typically deteriorates when extrapolation happens: namely, covariates shift to a region where the training distribution is scarce, and naturally, the learned model has little information. For robustness and regularization considerations, adversarial perturbation techniques are proposed as a remedy; however, more needs to be studied about what extrapolation region adversarial covariate shift will focus on, given a learned model. This paper precisely characterizes the extrapolation region, examining both regression and classification in an infinite-dimensional setting. We study the implications of adversarial covariate shifts to subsequent learning of the equilibrium -- the Bayes optimal model -- in a sequential game framework. We exploit the dynamics of the adversarial learning game and reveal the curious effects of the covariate shift to equilibrium learning and experimental design. In particular, we establish two directional convergence results that exhibit distinctive phenomena: (1) a blessing in regression, the adversarial covariate shifts in an exponential rate to an optimal experimental design for rapid subsequent learning, (2) a curse in classification, the adversarial covariate shifts in a subquadratic rate fast to the hardest experimental design trapping subsequent learning. ",
        "url": "https://arxiv.org/abs/2212.02457",
        "authors": [
          "Tengyuan Liang"
        ],
        "subjectives": [
          "Machine Learning (stat.ML)",
          "Machine Learning (cs.LG)",
          "Optimization and Control (math.OC)",
          "Statistics Theory (math.ST)"
        ]
      },
      {
        "id": "arXiv:2212.02477",
        "title": "Malaria Parasitic Detection using a New Deep Boosted and Ensemble  Learning Framework",
        "abstract": "Malaria is a potentially fatal plasmodium parasite injected by female anopheles mosquitoes that infect red blood cells and millions worldwide yearly. However, specialists' manual screening in clinical practice is laborious and prone to error. Therefore, a novel Deep Boosted and Ensemble Learning (DBEL) framework, comprising the stacking of new Boosted-BR-STM convolutional neural networks (CNN) and ensemble classifiers, is developed to screen malaria parasite images. The proposed STM-SB-BRNet is based on a new dilated-convolutional block-based split transform merge (STM) and feature-map Squeezing-Boosting (SB) ideas. Moreover, the new STM block uses regional and boundary operations to learn the malaria parasite's homogeneity, heterogeneity, and boundary with patterns. Furthermore, the diverse boosted channels are attained by employing Transfer Learning-based new feature-map SB in STM blocks at the abstract, medium, and conclusion levels to learn minute intensity and texture variation of the parasitic pattern. The proposed DBEL framework implicates the stacking of prominent and diverse boosted channels and provides the generated discriminative features of the developed Boosted-BR-STM to the ensemble of ML classifiers. The proposed framework improves the discrimination ability and generalization of ensemble learning. Moreover, the deep feature spaces of the developed Boosted-BR-STM and customized CNNs are fed into ML classifiers for comparative analysis. The proposed DBEL framework outperforms the existing techniques on the NIH malaria dataset that are enhanced using discrete wavelet transform to enrich feature space. The proposed DBEL framework achieved accuracy (98.50%), sensitivity (0.9920), F-score (0.9850), and AUC (0.997), which suggest it to be utilized for malaria parasite screening. ",
        "url": "https://arxiv.org/abs/2212.02477",
        "authors": [
          "Saddam Hussain Khan"
        ],
        "subjectives": [
          "Image and Video Processing (eess.IV)",
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:1805.10927",
        "title": "Scalable and Robust Community Detection with Randomized Sketching",
        "abstract": " Title: Scalable and Robust Community Detection with Randomized Sketching ",
        "url": "https://arxiv.org/abs/1805.10927",
        "authors": [
          "Mostafa Rahmani",
          "Andre Beckus",
          "Adel Karimian",
          "George Atia"
        ],
        "subjectives": [
          "Social and Information Networks (cs.SI)",
          "Machine Learning (cs.LG)",
          "Machine Learning (stat.ML)"
        ]
      },
      {
        "id": "arXiv:1906.06873",
        "title": "Running Time Analysis of the (1+1)-EA for Robust Linear Optimization",
        "abstract": " Comments: 17 pages, 1 table ",
        "url": "https://arxiv.org/abs/1906.06873",
        "authors": [
          "Chao Bian",
          "Chao Qian",
          "Ke Tang",
          "Yang Yu"
        ],
        "subjectives": [
          "Computational Complexity (cs.CC)",
          "Neural and Evolutionary Computing (cs.NE)",
          "Optimization and Control (math.OC)"
        ]
      },
      {
        "id": "arXiv:1912.13490",
        "title": "Representation Internal-Manipulation (RIM): A Neuro-Inspired  Computational Theory of Consciousness",
        "abstract": " Comments: 16 pages, 5 figures, preprint ",
        "url": "https://arxiv.org/abs/1912.13490",
        "authors": [
          "Gianluca Baldassarre",
          "Giovanni Granato"
        ],
        "subjectives": [
          "Artificial Intelligence (cs.AI)",
          "Machine Learning (cs.LG)",
          "Neural and Evolutionary Computing (cs.NE)",
          "Neurons and Cognition (q-bio.NC)",
          "Machine Learning (stat.ML)"
        ]
      },
      {
        "id": "arXiv:2001.08922",
        "title": "RePAD: Real-time Proactive Anomaly Detection for Time Series",
        "abstract": " Comments: 12 pages, 8 figures, the 34th International Conference on Advanced Information Networking and Applications (AINA 2020) ",
        "url": "https://arxiv.org/abs/2001.08922",
        "authors": [
          "Ming-Chang Lee",
          "Jia-Chun Lin",
          "Ernst Gunnar Gran"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Machine Learning (stat.ML)"
        ]
      },
      {
        "id": "arXiv:2004.02319",
        "title": "ReRe: A Lightweight Real-time Ready-to-Go Anomaly Detection Approach for  Time Series",
        "abstract": " Comments: 10 pages, 9 figures, COMPSAC 2020 ",
        "url": "https://arxiv.org/abs/2004.02319",
        "authors": [
          "Ming-Chang Lee",
          "Jia-Chun Lin",
          "Ernst Gunnar Gran"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Machine Learning (stat.ML)"
        ]
      },
      {
        "id": "arXiv:2004.06569",
        "title": "Improving Calibration and Out-of-Distribution Detection in Medical Image  Segmentation with Convolutional Neural Networks",
        "abstract": " Title: Improving Calibration and Out-of-Distribution Detection in Medical Image  Segmentation with Convolutional Neural Networks ",
        "url": "https://arxiv.org/abs/2004.06569",
        "authors": [
          "Davood Karimi",
          "Ali Gholipour"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Machine Learning (cs.LG)",
          "Image and Video Processing (eess.IV)",
          "Machine Learning (stat.ML)"
        ]
      },
      {
        "id": "arXiv:2009.02899",
        "title": "Quantifying Explainability of Saliency Methods in Deep Neural Networks  with a Synthetic Dataset",
        "abstract": " Title: Quantifying Explainability of Saliency Methods in Deep Neural Networks  with a Synthetic Dataset ",
        "url": "https://arxiv.org/abs/2009.02899",
        "authors": [
          "Erico Tjoa",
          "Cuntai Guan"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Artificial Intelligence (cs.AI)"
        ]
      },
      {
        "id": "arXiv:2009.11835",
        "title": "Sketch-based community detection in evolving networks",
        "abstract": " Title: Sketch-based community detection in evolving networks ",
        "url": "https://arxiv.org/abs/2009.11835",
        "authors": [
          "Andre Beckus",
          "George K. Atia"
        ],
        "subjectives": [
          "Physics and Society (physics.soc-ph)",
          "Machine Learning (cs.LG)",
          "Social and Information Networks (cs.SI)",
          "Machine Learning (stat.ML)"
        ]
      },
      {
        "id": "arXiv:2009.14115",
        "title": "CoKe: Localized Contrastive Learning for Robust Keypoint Detection",
        "abstract": " Comments: Accepted to WACV 2023 ",
        "url": "https://arxiv.org/abs/2009.14115",
        "authors": [
          "Yutong Bai",
          "Angtian Wang",
          "Adam Kortylewski",
          "Alan Yuille"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2103.10300",
        "title": "Error Prediction of Douglas-Rachford Algorithm for Linear Inverse  Problems: Asymptotics of Proximity Operator for Squared Loss",
        "abstract": " Comments: This work will be submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible ",
        "url": "https://arxiv.org/abs/2103.10300",
        "authors": [
          "Ryo Hayakawa"
        ],
        "subjectives": [
          "Signal Processing (eess.SP)",
          "Information Theory (cs.IT)",
          "Optimization and Control (math.OC)"
        ]
      },
      {
        "id": "arXiv:2105.04137",
        "title": "On the inversion number of oriented graphs",
        "abstract": " Title: On the inversion number of oriented graphs ",
        "url": "https://arxiv.org/abs/2105.04137",
        "authors": [
          "J\u00f8rgen Bang-Jensen",
          "Jonas Costa Ferreira da Silva",
          "Fr\u00e9d\u00e9ric Havet"
        ],
        "subjectives": [
          "Combinatorics (math.CO)",
          "Discrete Mathematics (cs.DM)"
        ]
      },
      {
        "id": "arXiv:2105.08269",
        "title": "Sparta: Spatially Attentive and Adversarially Robust Activation",
        "abstract": " Comments: 25 pages, 5 figures ",
        "url": "https://arxiv.org/abs/2105.08269",
        "authors": [
          "Qing Guo",
          "Felix Juefei-Xu",
          "Changqing Zhou",
          "Wei Feng",
          "Yang Liu",
          "Song Wang"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2106.15292",
        "title": "Adaptive Sample Selection for Robust Learning under Label Noise",
        "abstract": " Comments: Accepted at WACV 2023 ",
        "url": "https://arxiv.org/abs/2106.15292",
        "authors": [
          "Deep Patel",
          "P.S. Sastry"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Artificial Intelligence (cs.AI)",
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2107.04734",
        "title": "Layer-wise Analysis of a Self-supervised Speech Representation Model",
        "abstract": " Comments: Accepted to ASRU 2021. Code: this https URL ",
        "url": "https://arxiv.org/abs/2107.04734",
        "authors": [
          "Ankita Pasad",
          "Ju-Chieh Chou",
          "Karen Livescu"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)",
          "Machine Learning (cs.LG)",
          "Audio and Speech Processing (eess.AS)"
        ]
      },
      {
        "id": "arXiv:2108.12722",
        "title": "Feature Extraction for Machine Learning-based Intrusion Detection in IoT  Networks",
        "abstract": " Title: Feature Extraction for Machine Learning-based Intrusion Detection in IoT  Networks ",
        "url": "https://arxiv.org/abs/2108.12722",
        "authors": [
          "Mohanad Sarhan",
          "Siamak Layeghy",
          "Nour Moustafa",
          "Marcus Gallagher",
          "Marius Portmann"
        ],
        "subjectives": [
          "Networking and Internet Architecture (cs.NI)",
          "Cryptography and Security (cs.CR)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2110.01805",
        "title": "Self-Supervised Learning of Perceptually Optimized Block Motion  Estimates for Video Compression",
        "abstract": " Title: Self-Supervised Learning of Perceptually Optimized Block Motion  Estimates for Video Compression ",
        "url": "https://arxiv.org/abs/2110.01805",
        "authors": [
          "Somdyuti Paul",
          "Andrey Norkin",
          "Alan C. Bovik"
        ],
        "subjectives": [
          "Image and Video Processing (eess.IV)",
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2110.02929",
        "title": "Adversarial Attacks on Spiking Convolutional Neural Networks for  Event-based Vision",
        "abstract": " Comments: 9 pages plus Supplementary Material. Accepted in Frontiers in Neuroscience -- Neuromorphic Engineering ",
        "url": "https://arxiv.org/abs/2110.02929",
        "authors": [
          "Julian B\u00fcchel",
          "Gregor Lenz",
          "Yalun Hu",
          "Sadique Sheik",
          "Martino Sorbaro"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2110.04402",
        "title": "Walking into the complex plane to \"order\" better time integrators",
        "abstract": " Comments: 30 pages, 15 figures ",
        "url": "https://arxiv.org/abs/2110.04402",
        "authors": [
          "Jithin D. George",
          "Samuel Y. Jung",
          "Niall M. Mangan"
        ],
        "subjectives": [
          "Numerical Analysis (math.NA)",
          "Complex Variables (math.CV)"
        ]
      },
      {
        "id": "arXiv:2111.02630",
        "title": "Feature Learning and Network Structure from Noisy Node Activity Data",
        "abstract": " Title: Feature Learning and Network Structure from Noisy Node Activity Data ",
        "url": "https://arxiv.org/abs/2111.02630",
        "authors": [
          "Junyao Kuang",
          "Caterina Scoglio",
          "Kristin Michel"
        ],
        "subjectives": [
          "Networking and Internet Architecture (cs.NI)"
        ]
      },
      {
        "id": "arXiv:2111.03950",
        "title": "Kernel Methods for Multistage Causal Inference: Mediation Analysis and  Dynamic Treatment Effects",
        "abstract": " Comments: 88 pages. Material in this draft previously appeared in a working paper presented at the 2020 NeurIPS Workshop on ML for Economic Policy (arXiv:2010.04855v1). We have divided the original working paper (arXiv:2010.04855v1) into two projects: one paper focusing on static settings (arXiv:2010.04855) and this paper focusing on dynamic settings ",
        "url": "https://arxiv.org/abs/2111.03950",
        "authors": [
          "Rahul Singh",
          "Liyuan Xu",
          "Arthur Gretton"
        ],
        "subjectives": [
          "Methodology (stat.ME)",
          "Machine Learning (cs.LG)",
          "Econometrics (econ.EM)",
          "Machine Learning (stat.ML)"
        ]
      },
      {
        "id": "arXiv:2111.10613",
        "title": "Power Control in Cell-Free Massive MIMO Networks for UAVs URLLC under  the Finite Blocklength Regime",
        "abstract": " Comments: This paper has been accepted for publication on the IEEE Transactions on Communications. Only personal use of this material is permitted. Permission from IEEE must be obtained for all other uses ",
        "url": "https://arxiv.org/abs/2111.10613",
        "authors": [
          "Mohamed Elwekeil",
          "Alessio Zappone",
          "Stefano Buzzi"
        ],
        "subjectives": [
          "Information Theory (cs.IT)",
          "Signal Processing (eess.SP)"
        ]
      },
      {
        "id": "arXiv:2112.12591",
        "title": "Black-Box Testing of Deep Neural Networks through Test Case Diversity",
        "abstract": " Title: Black-Box Testing of Deep Neural Networks through Test Case Diversity ",
        "url": "https://arxiv.org/abs/2112.12591",
        "authors": [
          "Zohreh Aghababaeyan",
          "Manel Abdellatif",
          "Lionel Briand",
          "Ramesh S",
          "Mojtaba Bagherzadeh"
        ],
        "subjectives": [
          "Software Engineering (cs.SE)",
          "Artificial Intelligence (cs.AI)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2112.13267",
        "title": "Task and Model Agnostic Adversarial Attack on Graph Neural Networks",
        "abstract": " Comments: To appear as a full paper in AAAI 2023 ",
        "url": "https://arxiv.org/abs/2112.13267",
        "authors": [
          "Kartik Sharma",
          "Samidha Verma",
          "Sourav Medya",
          "Sayan Ranu",
          "Arnab Bhattacharya"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Cryptography and Security (cs.CR)"
        ]
      },
      {
        "id": "arXiv:2201.00150",
        "title": "Cross-Domain Deep Code Search with Few-Shot Meta Learning",
        "abstract": " Comments: Accepted by ICSE 2022 (The 44th International Conference on Software Engineering) ",
        "url": "https://arxiv.org/abs/2201.00150",
        "authors": [
          "Yitian Chai",
          "Hongyu Zhang",
          "Beijun Shen",
          "Xiaodong Gu"
        ],
        "subjectives": [
          "Software Engineering (cs.SE)"
        ]
      },
      {
        "id": "arXiv:2201.00965",
        "title": "Semantics-Preserved Distortion for Personal Privacy Protection in  Information Management",
        "abstract": " Title: Semantics-Preserved Distortion for Personal Privacy Protection in  Information Management ",
        "url": "https://arxiv.org/abs/2201.00965",
        "authors": [
          "Jiajia Li",
          "Letian Peng",
          "Ping Wang",
          "Zuchao Li",
          "Xueyi Li",
          "Hai Zhao"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)"
        ]
      },
      {
        "id": "arXiv:2201.01778",
        "title": "Quantum Capsule Networks",
        "abstract": " Comments: 7 pages (main text) + 8 pages (supplementary information), 8 figures ",
        "url": "https://arxiv.org/abs/2201.01778",
        "authors": [
          "Zidu Liu",
          "Pei-Xin Shen",
          "Weikang Li",
          "L.-M. Duan",
          "Dong-Ling Deng"
        ],
        "subjectives": [
          "Quantum Physics (quant-ph)",
          "Disordered Systems and Neural Networks (cond-mat.dis-nn)",
          "Mesoscale and Nanoscale Physics (cond-mat.mes-hall)",
          "Artificial Intelligence (cs.AI)",
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2201.10459",
        "title": "FRAMED: An AutoML Approach for Structural Performance Prediction of  Bicycle Frames",
        "abstract": " Title: FRAMED: An AutoML Approach for Structural Performance Prediction of  Bicycle Frames ",
        "url": "https://arxiv.org/abs/2201.10459",
        "authors": [
          "Lyle Regenwetter",
          "Colin Weaver",
          "Faez Ahmed"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Databases (cs.DB)"
        ]
      },
      {
        "id": "arXiv:2202.01315",
        "title": "Approximating Full Conformal Prediction at Scale via Influence Functions",
        "abstract": " Comments: 18 pages, 13 figures ",
        "url": "https://arxiv.org/abs/2202.01315",
        "authors": [
          "Javier Abad",
          "Umang Bhatt",
          "Adrian Weller",
          "Giovanni Cherubin"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Applications (stat.AP)"
        ]
      },
      {
        "id": "arXiv:2202.02669",
        "title": "SRPCN: Structure Retrieval based Point Completion Network",
        "abstract": " Comments: I think the proposed method has some defects ",
        "url": "https://arxiv.org/abs/2202.02669",
        "authors": [
          "Kaiyi Zhang",
          "Ximing Yang",
          "Yuan Wu",
          "Cheng Jin"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2202.08480",
        "title": "Eliciting Structural and Semantic Global Knowledge in Unsupervised Graph  Contrastive Learning",
        "abstract": " Comments: Accepted by AAAI2023 ",
        "url": "https://arxiv.org/abs/2202.08480",
        "authors": [
          "Kaize Ding",
          "Yancheng Wang",
          "Yingzhen Yang",
          "Huan Liu"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2202.09134",
        "title": "Quantifying the Effects of Data Augmentation",
        "abstract": " Title: Quantifying the Effects of Data Augmentation ",
        "url": "https://arxiv.org/abs/2202.09134",
        "authors": [
          "Kevin H. Huang",
          "Peter Orbanz",
          "Morgane Austern"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Statistics Theory (math.ST)",
          "Machine Learning (stat.ML)"
        ]
      },
      {
        "id": "arXiv:2202.11575",
        "title": "Shisha: Online scheduling of CNN pipelines on heterogeneous  architectures",
        "abstract": " Title: Shisha: Online scheduling of CNN pipelines on heterogeneous  architectures ",
        "url": "https://arxiv.org/abs/2202.11575",
        "authors": [
          "Pirah Noor Soomro",
          "Mustafa Abduljabbar",
          "Jeronimo Castrillon",
          "Miquel Peric\u00e0s"
        ],
        "subjectives": [
          "Performance (cs.PF)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2203.03099",
        "title": "Singular Value Perturbation and Deep Network Optimization",
        "abstract": " Comments: Constr Approx (2022) ",
        "url": "https://arxiv.org/abs/2203.03099",
        "authors": [
          "Rudolf H. Riedi",
          "Randall Balestriero",
          "Richard G. Baraniuk"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Spectral Theory (math.SP)"
        ]
      },
      {
        "id": "arXiv:2203.04300",
        "title": "Multi-trial Neural Architecture Search with Lottery Tickets",
        "abstract": " Title: Multi-trial Neural Architecture Search with Lottery Tickets ",
        "url": "https://arxiv.org/abs/2203.04300",
        "authors": [
          "Zimian Wei",
          "Hengyue Pan",
          "Lujun Li",
          "Menglong Lu",
          "Xin Niu",
          "Peijie Dong",
          "Dongsheng Li"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Neural and Evolutionary Computing (cs.NE)"
        ]
      },
      {
        "id": "arXiv:2203.15793",
        "title": "Instance Relation Graph Guided Source-Free Domain Adaptive Object  Detection",
        "abstract": " Title: Instance Relation Graph Guided Source-Free Domain Adaptive Object  Detection ",
        "url": "https://arxiv.org/abs/2203.15793",
        "authors": [
          "Vibashan VS",
          "Poojan Oza",
          "Vishal M. Patel"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2204.01054",
        "title": "Combinatorial refinement on circulant graphs",
        "abstract": " Comments: 20 pages, 1 figure ",
        "url": "https://arxiv.org/abs/2204.01054",
        "authors": [
          "Laurence Kluge"
        ],
        "subjectives": [
          "Combinatorics (math.CO)",
          "Computational Complexity (cs.CC)",
          "Discrete Mathematics (cs.DM)"
        ]
      },
      {
        "id": "arXiv:2204.06173",
        "title": "Synthesizing Adversarial Visual Scenarios for Model-Based Robotic  Control",
        "abstract": " Comments: Conference on Robot Learning, 2022 ",
        "url": "https://arxiv.org/abs/2204.06173",
        "authors": [
          "Shubhankar Agarwal",
          "Sandeep P. Chinchali"
        ],
        "subjectives": [
          "Robotics (cs.RO)"
        ]
      },
      {
        "id": "arXiv:2204.07629",
        "title": "Navigation between initial and desired community states using shortcuts",
        "abstract": " Title: Navigation between initial and desired community states using shortcuts ",
        "url": "https://arxiv.org/abs/2204.07629",
        "authors": [
          "Benjamin W. Blonder",
          "Michael H. Lim",
          "Zachary Sunberg",
          "Claire Tomlin"
        ],
        "subjectives": [
          "Populations and Evolution (q-bio.PE)",
          "Systems and Control (eess.SY)"
        ]
      },
      {
        "id": "arXiv:2204.09710",
        "title": "Complete identification of complex salt geometries from inaccurate  migrated subsurface offset gathers using deep learning",
        "abstract": " Comments: Manuscript published at Geophysics ",
        "url": "https://arxiv.org/abs/2204.09710",
        "authors": [
          "Ana Paula O. Muller",
          "Jesse C. Costa",
          "Clecio R. Bom",
          "Elisangela L. Faria",
          "Matheus Klatt",
          "Gabriel Teixeira",
          "Marcelo P. de Albuquerque",
          "Marcio P. de Albuquerque"
        ],
        "subjectives": [
          "Geophysics (physics.geo-ph)",
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2204.11795",
        "title": "Performer: A Novel PPG-to-ECG Reconstruction Transformer for a Digital  Biomarker of Cardiovascular Disease Detection",
        "abstract": " Comments: 8 pages, 11 figures ",
        "url": "https://arxiv.org/abs/2204.11795",
        "authors": [
          "Ella Lan"
        ],
        "subjectives": [
          "Signal Processing (eess.SP)",
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Machine Learning (cs.LG)",
          "Image and Video Processing (eess.IV)"
        ]
      },
      {
        "id": "arXiv:2204.13257",
        "title": "Joint User Association and Beamforming in Integrated  Satellite-HAPS-Ground Networks",
        "abstract": " Title: Joint User Association and Beamforming in Integrated  Satellite-HAPS-Ground Networks ",
        "url": "https://arxiv.org/abs/2204.13257",
        "authors": [
          "Shasha Liu",
          "Hayssam Dahrouj",
          "Mohamed-Slim Alouini"
        ],
        "subjectives": [
          "Information Theory (cs.IT)",
          "Networking and Internet Architecture (cs.NI)"
        ]
      },
      {
        "id": "arXiv:2205.03976",
        "title": "Orientations and cycles in supersingular isogeny graphs",
        "abstract": " Comments: 41 pages, 7 figures ",
        "url": "https://arxiv.org/abs/2205.03976",
        "authors": [
          "Sarah Arpin",
          "Mingjie Chen",
          "Kristin E. Lauter",
          "Renate Scheidler",
          "Katherine E. Stange",
          "Ha T. N. Tran"
        ],
        "subjectives": [
          "Number Theory (math.NT)",
          "Cryptography and Security (cs.CR)"
        ]
      },
      {
        "id": "arXiv:2205.12606",
        "title": "ReSmooth: Detecting and Utilizing OOD Samples when Training with Data  Augmentation",
        "abstract": " Comments: The paper is accepted as a TNNLS regular paper. See the published version in \"Early Access\" area on IEEE Xplore: this https URL ",
        "url": "https://arxiv.org/abs/2205.12606",
        "authors": [
          "Chenyang Wang",
          "Junjun Jiang",
          "Xiong Zhou",
          "Xianming Liu"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2205.13148",
        "title": "Grammar Detection for Sentiment Analysis through Improved Viterbi  Algorithm",
        "abstract": " Comments: Updating the current work ",
        "url": "https://arxiv.org/abs/2205.13148",
        "authors": [
          "Surya Teja Chavali",
          "Charan Tej Kandavalli",
          "Sugash T M"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2205.13958",
        "title": "Machine Learning-Based User Scheduling in Integrated  Satellite-HAPS-Ground Networks",
        "abstract": " Comments: arXiv admin note: substantial text overlap with arXiv:2204.13257 ",
        "url": "https://arxiv.org/abs/2205.13958",
        "authors": [
          "Hayssam Dahrouj",
          "Shasha Liu",
          "Mohamed-Slim Alouini"
        ],
        "subjectives": [
          "Artificial Intelligence (cs.AI)"
        ]
      },
      {
        "id": "arXiv:2205.14690",
        "title": "CoNT: Contrastive Neural Text Generation",
        "abstract": " Comments: Accepted by NeurIPS 2022 ",
        "url": "https://arxiv.org/abs/2205.14690",
        "authors": [
          "Chenxin An",
          "Jiangtao Feng",
          "Kai Lv",
          "Lingpeng Kong",
          "Xipeng Qiu",
          "Xuanjing Huang"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)"
        ]
      },
      {
        "id": "arXiv:2205.15128",
        "title": "Domain Constraints in Feature Space: Strengthening Robustness of Android  Malware Detection against Realizable Adversarial Examples",
        "abstract": " Title: Domain Constraints in Feature Space: Strengthening Robustness of Android  Malware Detection against Realizable Adversarial Examples ",
        "url": "https://arxiv.org/abs/2205.15128",
        "authors": [
          "Hamid Bostani",
          "Zhuoran Liu",
          "Zhengyu Zhao",
          "Veelasha Moonsamy"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Cryptography and Security (cs.CR)"
        ]
      },
      {
        "id": "arXiv:2206.00383",
        "title": "Neural Improvement Heuristics for Graph Combinatorial Optimization  Problems",
        "abstract": " Title: Neural Improvement Heuristics for Graph Combinatorial Optimization  Problems ",
        "url": "https://arxiv.org/abs/2206.00383",
        "authors": [
          "Andoni I. Garmendia",
          "Josu Ceberio",
          "Alexander Mendiburu"
        ],
        "subjectives": [
          "Artificial Intelligence (cs.AI)",
          "Discrete Mathematics (cs.DM)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2206.01995",
        "title": "Combinatorial Causal Bandits",
        "abstract": " Comments: 30 pages, 9 figures ",
        "url": "https://arxiv.org/abs/2206.01995",
        "authors": [
          "Shi Feng",
          "Wei Chen"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Social and Information Networks (cs.SI)",
          "Methodology (stat.ME)",
          "Machine Learning (stat.ML)"
        ]
      },
      {
        "id": "arXiv:2206.02027",
        "title": "Implicit Neural Representation for Mesh-Free Inverse Obstacle Scattering",
        "abstract": " Comments: 6 pages, 8 figures, to be published in 2022 Asilomar Conference on Signals, Systems, and Computers ",
        "url": "https://arxiv.org/abs/2206.02027",
        "authors": [
          "Tin Vla\u0161i\u0107",
          "Hieu Nguyen",
          "AmirEhsan Khorashadizadeh",
          "Ivan Dokmani\u0107"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Signal Processing (eess.SP)"
        ]
      },
      {
        "id": "arXiv:2206.05904",
        "title": "Superiority of GNN over NN in generalizing bandlimited functions",
        "abstract": " Title: Superiority of GNN over NN in generalizing bandlimited functions ",
        "url": "https://arxiv.org/abs/2206.05904",
        "authors": [
          "A. Martina Neuman",
          "Rongrong Wang",
          "Yuying Xie"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Machine Learning (stat.ML)"
        ]
      },
      {
        "id": "arXiv:2206.06757",
        "title": "RoSGAS: Adaptive Social Bot Detection with Reinforced Self-Supervised  GNN Architecture Search",
        "abstract": " Comments: 32 pages, 12 figures accpted by ACM Transactions on the Web (TWEB) ",
        "url": "https://arxiv.org/abs/2206.06757",
        "authors": [
          "Yingguang Yang",
          "Renyu Yang",
          "Yangyang Li",
          "Kai Cui",
          "Zhiqin Yang",
          "Yue Wang",
          "Jie Xu",
          "Haiyong Xie"
        ],
        "subjectives": [
          "Social and Information Networks (cs.SI)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2206.07758",
        "title": "Reconstructing Training Data from Trained Neural Networks",
        "abstract": " Comments: Fixed a typo in the acknowledgements ",
        "url": "https://arxiv.org/abs/2206.07758",
        "authors": [
          "Niv Haim",
          "Gal Vardi",
          "Gilad Yehudai",
          "Ohad Shamir",
          "Michal Irani"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Cryptography and Security (cs.CR)",
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Neural and Evolutionary Computing (cs.NE)",
          "Machine Learning (stat.ML)"
        ]
      },
      {
        "id": "arXiv:2206.10305",
        "title": "Scale-Variant Robust Kernel Optimization for Non-linear Least Squares  Problems",
        "abstract": " Comments: Submitted to IEEE Transactions on Aerospace and Electronic Systems ",
        "url": "https://arxiv.org/abs/2206.10305",
        "authors": [
          "Shounak Das",
          "Jason Gross"
        ],
        "subjectives": [
          "Robotics (cs.RO)"
        ]
      },
      {
        "id": "arXiv:2206.10911",
        "title": "Influence of uncertainty estimation techniques on false-positive  reduction in liver lesion detection",
        "abstract": " Comments: Submitted to The Journal of Machine Learning for Biomedical Imaging ",
        "url": "https://arxiv.org/abs/2206.10911",
        "authors": [
          "Ishaan Bhat",
          "Josien P.W. Pluim",
          "Max A. Viergever",
          "Hugo J. Kuijf"
        ],
        "subjectives": [
          "Image and Video Processing (eess.IV)",
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2206.11241",
        "title": "Concentration inequalities and optimal number of layers for stochastic  deep neural networks",
        "abstract": " Title: Concentration inequalities and optimal number of layers for stochastic  deep neural networks ",
        "url": "https://arxiv.org/abs/2206.11241",
        "authors": [
          "Michele Caprio",
          "Sayan Mukherjee"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Machine Learning (stat.ML)"
        ]
      },
      {
        "id": "arXiv:2206.12344",
        "title": "Segmentation-free PVC for Cardiac SPECT using a Densely-connected  Multi-dimensional Dynamic Network",
        "abstract": " Comments: 12 pages, 11 figures. Accepted for publication at IEEE Transactions on Medical Imaging ",
        "url": "https://arxiv.org/abs/2206.12344",
        "authors": [
          "Huidong Xie",
          "Zhao Liu",
          "Luyao Shi",
          "Kathleen Greco",
          "Xiongchao Chen",
          "Bo Zhou",
          "Attila Feher",
          "John C. Stendahl",
          "Nabil Boutagy",
          "Tassos C. Kyriakides",
          "Ge Wang",
          "Albert J. Sinusas",
          "Chi Liu"
        ],
        "subjectives": [
          "Image and Video Processing (eess.IV)",
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2207.02547",
        "title": "Simple and Efficient Heterogeneous Graph Neural Network",
        "abstract": " Comments: To appear at AAAI 2023 ",
        "url": "https://arxiv.org/abs/2207.02547",
        "authors": [
          "Xiaocheng Yang",
          "Mingyu Yan",
          "Shirui Pan",
          "Xiaochun Ye",
          "Dongrui Fan"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2207.02614",
        "title": "Securing Optimized Code Against Power Side Channels",
        "abstract": " Title: Securing Optimized Code Against Power Side Channels ",
        "url": "https://arxiv.org/abs/2207.02614",
        "authors": [
          "Rodothea Myrsini Tsoupidi",
          "Roberto Casta\u00f1eda Lozano",
          "Elena Troubitsyna",
          "Panagiotis Papadimitratos"
        ],
        "subjectives": [
          "Cryptography and Security (cs.CR)",
          "Artificial Intelligence (cs.AI)"
        ]
      },
      {
        "id": "arXiv:2207.03020",
        "title": "Mechanisms of True and False Rumor Sharing in Social Media: Collective  Intelligence or Herd Behavior?",
        "abstract": " Title: Mechanisms of True and False Rumor Sharing in Social Media: Collective  Intelligence or Herd Behavior? ",
        "url": "https://arxiv.org/abs/2207.03020",
        "authors": [
          "Nicolas Pr\u00f6llochs",
          "Stefan Feuerriegel"
        ],
        "subjectives": [
          "Social and Information Networks (cs.SI)"
        ]
      },
      {
        "id": "arXiv:2207.03997",
        "title": "A Tutorial on Event Detection using Social Media Data Analysis:  Applications, Challenges, and Open Problems",
        "abstract": " Title: A Tutorial on Event Detection using Social Media Data Analysis:  Applications, Challenges, and Open Problems ",
        "url": "https://arxiv.org/abs/2207.03997",
        "authors": [
          "Mohammadsepehr Karimiziarani"
        ],
        "subjectives": [
          "Social and Information Networks (cs.SI)"
        ]
      },
      {
        "id": "arXiv:2208.05740",
        "title": "General Cutting Planes for Bound-Propagation-Based Neural Network  Verification",
        "abstract": " Comments: Accepted by NeurIPS 2022. GCP-CROWN is part of the alpha-beta-CROWN verifier, the VNN-COMP 2022 winner ",
        "url": "https://arxiv.org/abs/2208.05740",
        "authors": [
          "Huan Zhang",
          "Shiqi Wang",
          "Kaidi Xu",
          "Linyi Li",
          "Bo Li",
          "Suman Jana",
          "Cho-Jui Hsieh",
          "J. Zico Kolter"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Cryptography and Security (cs.CR)",
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Optimization and Control (math.OC)",
          "Machine Learning (stat.ML)"
        ]
      },
      {
        "id": "arXiv:2208.06073",
        "title": "Conditional Antibody Design as 3D Equivariant Graph Translation",
        "abstract": " Comments: Preprint. Under review ",
        "url": "https://arxiv.org/abs/2208.06073",
        "authors": [
          "Xiangzhe Kong",
          "Wenbing Huang",
          "Yang Liu"
        ],
        "subjectives": [
          "Biomolecules (q-bio.BM)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2208.09446",
        "title": "MonoSIM: Simulating Learning Behaviors of Heterogeneous Point Cloud  Object Detectors for Monocular 3D Object Detection",
        "abstract": " Title: MonoSIM: Simulating Learning Behaviors of Heterogeneous Point Cloud  Object Detectors for Monocular 3D Object Detection ",
        "url": "https://arxiv.org/abs/2208.09446",
        "authors": [
          "Han Sun",
          "Zhaoxin Fan",
          "Zhenbo Song",
          "Zhicheng Wang",
          "Kejian Wu",
          "Jianfeng Lu"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2208.10741",
        "title": "Hierarchically Decomposed Graph Convolutional Networks for  Skeleton-Based Action Recognition",
        "abstract": " Comments: 12 pages, 7 figures ",
        "url": "https://arxiv.org/abs/2208.10741",
        "authors": [
          "Jungho Lee",
          "Minhyeok Lee",
          "Dogyoon Lee",
          "Sangyoun Lee"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2209.06395",
        "title": "Point Cloud Registration-Driven Robust Feature Matching for 3D Siamese  Object Tracking",
        "abstract": " Title: Point Cloud Registration-Driven Robust Feature Matching for 3D Siamese  Object Tracking ",
        "url": "https://arxiv.org/abs/2209.06395",
        "authors": [
          "Haobo Jiang",
          "Kaihao Lan",
          "Le Hui",
          "Guangyu Li",
          "Jin Xie",
          "Jian Yang"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Artificial Intelligence (cs.AI)"
        ]
      },
      {
        "id": "arXiv:2209.06932",
        "title": "Optimizing Connectivity through Network Gradients for Restricted  Boltzmann Machines",
        "abstract": " Title: Optimizing Connectivity through Network Gradients for Restricted  Boltzmann Machines ",
        "url": "https://arxiv.org/abs/2209.06932",
        "authors": [
          "A. C. N. de Oliveira",
          "D. R. Figueiredo"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2209.08932",
        "title": "OPR-Miner: Order-preserving rule mining for time series",
        "abstract": " Title: OPR-Miner: Order-preserving rule mining for time series ",
        "url": "https://arxiv.org/abs/2209.08932",
        "authors": [
          "Youxi Wu",
          "Xiaoqian Zhao",
          "Yan Li",
          "Lei Guo",
          "Xingquan Zhu",
          "Philippe Fournier-Viger",
          "Xindong Wu"
        ],
        "subjectives": [
          "Databases (cs.DB)"
        ]
      },
      {
        "id": "arXiv:2209.10100",
        "title": "Flashlight: Scalable Link Prediction with Effective Decoders",
        "abstract": " Comments: arXiv admin note: text overlap with arXiv:2112.02936 by other authors ",
        "url": "https://arxiv.org/abs/2209.10100",
        "authors": [
          "Yiwei Wang",
          "Bryan Hooi",
          "Yozen Liu",
          "Tong Zhao",
          "Zhichun Guo",
          "Neil Shah"
        ],
        "subjectives": [
          "Social and Information Networks (cs.SI)",
          "Artificial Intelligence (cs.AI)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2209.11142",
        "title": "A Generalist Neural Algorithmic Learner",
        "abstract": " Comments: To appear at LoG 2022 (Spotlight talk). 23 pages, 11 figures ",
        "url": "https://arxiv.org/abs/2209.11142",
        "authors": [
          "Borja Ibarz",
          "Vitaly Kurin",
          "George Papamakarios",
          "Kyriacos Nikiforou",
          "Mehdi Bennani",
          "R\u00f3bert Csord\u00e1s",
          "Andrew Dudzik",
          "Matko Bo\u0161njak",
          "Alex Vitvitskyi",
          "Yulia Rubanova",
          "Andreea Deac",
          "Beatrice Bevilacqua",
          "Yaroslav Ganin",
          "Charles Blundell",
          "Petar Veli\u010dkovi\u0107"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Artificial Intelligence (cs.AI)",
          "Machine Learning (stat.ML)"
        ]
      },
      {
        "id": "arXiv:2209.11388",
        "title": "LGDN: Language-Guided Denoising Network for Video-Language Modeling",
        "abstract": " Comments: Accepted by NeurIPS2022 ",
        "url": "https://arxiv.org/abs/2209.11388",
        "authors": [
          "Haoyu Lu",
          "Mingyu Ding",
          "Nanyi Fei",
          "Yuqi Huo",
          "Zhiwu Lu"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Artificial Intelligence (cs.AI)",
          "Multimedia (cs.MM)"
        ]
      },
      {
        "id": "arXiv:2209.12457",
        "title": "Fault Detection Scheme for Grid-Forming Inverters in Islanded  Droop-Controlled AC Microgrids",
        "abstract": " Comments: 10 pages, 7 figures ",
        "url": "https://arxiv.org/abs/2209.12457",
        "authors": [
          "Gabriel Intriago",
          "Andres Intriago",
          "Raul Intriago",
          "Yu Zhang"
        ],
        "subjectives": [
          "Systems and Control (eess.SY)"
        ]
      },
      {
        "id": "arXiv:2210.01425",
        "title": "Unveiling the Black Box of PLMs with Semantic Anchors: Towards  Interpretable Neural Semantic Parsing",
        "abstract": " Comments: AAAI 2023 Main Track Long Paper ",
        "url": "https://arxiv.org/abs/2210.01425",
        "authors": [
          "Lunyiu Nie",
          "Jiuding Sun",
          "Yanlin Wang",
          "Lun Du",
          "Lei Hou",
          "Juanzi Li",
          "Shi Han",
          "Dongmei Zhang",
          "Jidong Zhai"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)",
          "Artificial Intelligence (cs.AI)"
        ]
      },
      {
        "id": "arXiv:2210.01549",
        "title": "Diffusion Models for Graphs Benefit From Discrete State Spaces",
        "abstract": " Comments: Presented at the First Learning on Graphs Conference (LoG 2022) and the NeurIPS 2022 New Frontiers in Graph Learning Workshop (NeurIPS GLFrontiers 2022) ",
        "url": "https://arxiv.org/abs/2210.01549",
        "authors": [
          "Kilian Konstantin Haefeli",
          "Karolis Martinkus",
          "Nathana\u00ebl Perraudin",
          "Roger Wattenhofer"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Social and Information Networks (cs.SI)",
          "Machine Learning (stat.ML)"
        ]
      },
      {
        "id": "arXiv:2210.08388",
        "title": "RoS-KD: A Robust Stochastic Knowledge Distillation Approach for Noisy  Medical Imaging",
        "abstract": " Comments: Accepted in ICDM 2022 ",
        "url": "https://arxiv.org/abs/2210.08388",
        "authors": [
          "Ajay Jaiswal",
          "Kumar Ashutosh",
          "Justin F Rousseau",
          "Yifan Peng",
          "Zhangyang Wang",
          "Ying Ding"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2210.12438",
        "title": "Algorithms with Prediction Portfolios",
        "abstract": " Comments: 24 pages. Appears at NeurIPS 2022 ",
        "url": "https://arxiv.org/abs/2210.12438",
        "authors": [
          "Michael Dinitz",
          "Sungjin Im",
          "Thomas Lavastida",
          "Benjamin Moseley",
          "Sergei Vassilvitskii"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Data Structures and Algorithms (cs.DS)"
        ]
      },
      {
        "id": "arXiv:2210.15858",
        "title": "Vox-Fusion: Dense Tracking and Mapping with Voxel-based Neural Implicit  Representation",
        "abstract": " Title: Vox-Fusion: Dense Tracking and Mapping with Voxel-based Neural Implicit  Representation ",
        "url": "https://arxiv.org/abs/2210.15858",
        "authors": [
          "Xingrui Yang",
          "Hai Li",
          "Hongjia Zhai",
          "Yuhang Ming",
          "Yuqian Liu",
          "Guofeng Zhang"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Graphics (cs.GR)",
          "Robotics (cs.RO)"
        ]
      },
      {
        "id": "arXiv:2211.03929",
        "title": "Comparative layer-wise analysis of self-supervised speech models",
        "abstract": " Comments: Submitted to ICASSP 2022. Code: this https URL ",
        "url": "https://arxiv.org/abs/2211.03929",
        "authors": [
          "Ankita Pasad",
          "Bowen Shi",
          "Karen Livescu"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)",
          "Machine Learning (cs.LG)",
          "Sound (cs.SD)",
          "Audio and Speech Processing (eess.AS)"
        ]
      },
      {
        "id": "arXiv:2211.06352",
        "title": "Spectral Triadic Decompositions of Real-World Networks",
        "abstract": " Title: Spectral Triadic Decompositions of Real-World Networks ",
        "url": "https://arxiv.org/abs/2211.06352",
        "authors": [
          "Sabyasachi Basu",
          "Suman Kalyan Bera",
          "C. Seshadhri"
        ],
        "subjectives": [
          "Social and Information Networks (cs.SI)",
          "Discrete Mathematics (cs.DM)",
          "Data Structures and Algorithms (cs.DS)"
        ]
      },
      {
        "id": "arXiv:2211.06482",
        "title": "Augmenting Transformer-Transducer Based Speaker Change Detection With  Token-Level Training Loss",
        "abstract": " Title: Augmenting Transformer-Transducer Based Speaker Change Detection With  Token-Level Training Loss ",
        "url": "https://arxiv.org/abs/2211.06482",
        "authors": [
          "Guanlong Zhao",
          "Quan Wang",
          "Han Lu",
          "Yiling Huang",
          "Ignacio Lopez Moreno"
        ],
        "subjectives": [
          "Audio and Speech Processing (eess.AS)",
          "Machine Learning (cs.LG)",
          "Sound (cs.SD)"
        ]
      },
      {
        "id": "arXiv:2211.06548",
        "title": "Computationally Light Spectrally Normalized Memory Neuron Network based  Estimator for GPS-Denied operation of Micro UAV",
        "abstract": " Comments: Submitted to L4DC 2023 ",
        "url": "https://arxiv.org/abs/2211.06548",
        "authors": [
          "Nishanth Rao",
          "Suresh Sundaram",
          "Varun Raghavendra"
        ],
        "subjectives": [
          "Robotics (cs.RO)"
        ]
      },
      {
        "id": "arXiv:2211.06627",
        "title": "MARLIN: Masked Autoencoder for facial video Representation LearnINg",
        "abstract": " Title: MARLIN: Masked Autoencoder for facial video Representation LearnINg ",
        "url": "https://arxiv.org/abs/2211.06627",
        "authors": [
          "Zhixi Cai",
          "Shreya Ghosh",
          "Kalin Stefanov",
          "Abhinav Dhall",
          "Jianfei Cai",
          "Hamid Rezatofighi",
          "Reza Haffari",
          "Munawar Hayat"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2211.07636",
        "title": "EVA: Exploring the Limits of Masked Visual Representation Learning at  Scale",
        "abstract": " Comments: v2: (i) fix / update EVA IN-1K variants results. (ii) add / update EVA-CLIP results. (iii) add Appendix. (iv) release all the code and models at this https URL ",
        "url": "https://arxiv.org/abs/2211.07636",
        "authors": [
          "Yuxin Fang",
          "Wen Wang",
          "Binhui Xie",
          "Quan Sun",
          "Ledell Wu",
          "Xinggang Wang",
          "Tiejun Huang",
          "Xinlong Wang",
          "Yue Cao"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Computation and Language (cs.CL)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2211.11907",
        "title": "Robust Faber--Schauder approximation based on discrete observations of  an antiderivative",
        "abstract": " Title: Robust Faber--Schauder approximation based on discrete observations of  an antiderivative ",
        "url": "https://arxiv.org/abs/2211.11907",
        "authors": [
          "Xiyue Han",
          "Alexander Schied"
        ],
        "subjectives": [
          "Numerical Analysis (math.NA)",
          "Statistical Finance (q-fin.ST)"
        ]
      },
      {
        "id": "arXiv:2211.12933",
        "title": "Join the High Accuracy Club on ImageNet with A Binary Neural Network  Ticket",
        "abstract": " Title: Join the High Accuracy Club on ImageNet with A Binary Neural Network  Ticket ",
        "url": "https://arxiv.org/abs/2211.12933",
        "authors": [
          "Nianhui Guo",
          "Joseph Bethge",
          "Christoph Meinel",
          "Haojin Yang"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Artificial Intelligence (cs.AI)"
        ]
      },
      {
        "id": "arXiv:2211.13231",
        "title": "Predicting Biomedical Interactions with Probabilistic Model Selection  for Graph Neural Networks",
        "abstract": " Title: Predicting Biomedical Interactions with Probabilistic Model Selection  for Graph Neural Networks ",
        "url": "https://arxiv.org/abs/2211.13231",
        "authors": [
          "Kishan KC",
          "Rui Li",
          "Paribesh Regmi",
          "Anne R. Haake"
        ],
        "subjectives": [
          "Quantitative Methods (q-bio.QM)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2211.13305",
        "title": "Dual Graphs of Polyhedral Decompositions for the Detection of  Adversarial Attacks",
        "abstract": " Comments: 978-1-6654-8045-1/22/\\$31.00 \\copyright{}2022 IEEE The 6th Workshop on Graph Techniques for Adversarial Activity Analytics (GTA 2022) ",
        "url": "https://arxiv.org/abs/2211.13305",
        "authors": [
          "Huma Jamil",
          "Yajing Liu",
          "Christina M. Cole",
          "Nathaniel Blanchard",
          "Emily J. King",
          "Michael Kirby",
          "Christopher Peterson"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Cryptography and Security (cs.CR)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2211.13327",
        "title": "A Report on the Euphemisms Detection Shared Task",
        "abstract": " Title: A Report on the Euphemisms Detection Shared Task ",
        "url": "https://arxiv.org/abs/2211.13327",
        "authors": [
          "Patrick Lee",
          "Anna Feldman",
          "Jing Peng"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)",
          "Artificial Intelligence (cs.AI)"
        ]
      },
      {
        "id": "arXiv:2211.13897",
        "title": "AFR-Net: Attention-Driven Fingerprint Recognition Network",
        "abstract": " Title: AFR-Net: Attention-Driven Fingerprint Recognition Network ",
        "url": "https://arxiv.org/abs/2211.13897",
        "authors": [
          "Steven A. Grosz",
          "Anil K. Jain"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2211.14489",
        "title": "Mitigating Relational Bias on Knowledge Graphs",
        "abstract": " Title: Mitigating Relational Bias on Knowledge Graphs ",
        "url": "https://arxiv.org/abs/2211.14489",
        "authors": [
          "Yu-Neng Chuang",
          "Kwei-Herng Lai",
          "Ruixiang Tang",
          "Mengnan Du",
          "Chia-Yuan Chang",
          "Na Zou",
          "Xia Hu"
        ],
        "subjectives": [
          "Artificial Intelligence (cs.AI)",
          "Computers and Society (cs.CY)",
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2211.15030",
        "title": "Imperceptible Adversarial Attack via Invertible Neural Networks",
        "abstract": " Title: Imperceptible Adversarial Attack via Invertible Neural Networks ",
        "url": "https://arxiv.org/abs/2211.15030",
        "authors": [
          "Zihan Chen",
          "Ziyue Wang",
          "Junjie Huang",
          "Wentao Zhao",
          "Xiao Liu",
          "Dejian Guan"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)",
          "Image and Video Processing (eess.IV)"
        ]
      },
      {
        "id": "arXiv:2211.15223",
        "title": "Gamma-convergence of a nonlocal perimeter arising in adversarial machine  learning",
        "abstract": " Comments: Improved convergence results for adversarial training ",
        "url": "https://arxiv.org/abs/2211.15223",
        "authors": [
          "Leon Bungert",
          "Kerrek Stinson"
        ],
        "subjectives": [
          "Analysis of PDEs (math.AP)",
          "Machine Learning (cs.LG)",
          "Optimization and Control (math.OC)"
        ]
      },
      {
        "id": "arXiv:2211.15279",
        "title": "Establishment of Neural Networks Robust to Label Noise",
        "abstract": " Comments: 11 pages, 7 figures ",
        "url": "https://arxiv.org/abs/2211.15279",
        "authors": [
          "Pengwei Yang",
          "Chongyangzi Teng",
          "Jack George Mangos"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)"
        ]
      },
      {
        "id": "arXiv:2211.15899",
        "title": "FakeEdge: Alleviate Dataset Shift in Link Prediction",
        "abstract": " Comments: Accepted to Learning on Graph ",
        "url": "https://arxiv.org/abs/2211.15899",
        "authors": [
          "Kaiwen Dong",
          "Yijun Tian",
          "Zhichun Guo",
          "Yang Yang",
          "Nitesh V. Chawla"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Social and Information Networks (cs.SI)",
          "Machine Learning (stat.ML)"
        ]
      },
      {
        "id": "arXiv:2211.15980",
        "title": "End-to-End Neural Discourse Deixis Resolution in Dialogue",
        "abstract": " Comments: Accepted as a long paper to EMNLP 2022 ",
        "url": "https://arxiv.org/abs/2211.15980",
        "authors": [
          "Shengjie Li",
          "Vincent Ng"
        ],
        "subjectives": [
          "Computation and Language (cs.CL)"
        ]
      },
      {
        "id": "arXiv:2211.16235",
        "title": "DCDetector: An IoT terminal vulnerability mining system based on  distributed deep ensemble learning under source code representation",
        "abstract": " Comments: Some experiments need to be done better, and some theories need to be improved,thank you ",
        "url": "https://arxiv.org/abs/2211.16235",
        "authors": [
          "Wen Zhou"
        ],
        "subjectives": [
          "Cryptography and Security (cs.CR)",
          "Artificial Intelligence (cs.AI)"
        ]
      },
      {
        "id": "arXiv:2212.00133",
        "title": "{Generative Adversarial Learning of Sinkhorn Algorithm Initializations}",
        "abstract": " Comments: Added universality result for the generator. Improved literature review. Fixed typos. 11 pages, 7 figures ",
        "url": "https://arxiv.org/abs/2212.00133",
        "authors": [
          "Jonathan Geuter",
          "Vaios Laschos"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Optimization and Control (math.OC)",
          "Machine Learning (stat.ML)"
        ]
      },
      {
        "id": "arXiv:2212.00554",
        "title": "Early prediction of the risk of ICU mortality with Deep Federated  Learning",
        "abstract": " Comments: 12 pages ",
        "url": "https://arxiv.org/abs/2212.00554",
        "authors": [
          "Korbinian Randl",
          "N\u00faria Llad\u00f3s Armengol",
          "Lena Mondrejevski",
          "Ioanna Miliou"
        ],
        "subjectives": [
          "Machine Learning (cs.LG)",
          "Artificial Intelligence (cs.AI)"
        ]
      },
      {
        "id": "arXiv:2212.00565",
        "title": "Weakly-supervised detection of AMD-related lesions in color fundus  images using explainable deep learning",
        "abstract": " Comments: Accepted in the journal Computer Methods and Programs in Biomedicine on November 29, 2022 ",
        "url": "https://arxiv.org/abs/2212.00565",
        "authors": [
          "Jos\u00e9 Morano",
          "\u00c1lvaro S. Hervella",
          "Jos\u00e9 Rouco",
          "Jorge Novo",
          "Jos\u00e9 I. Fern\u00e1ndez-Vigo",
          "Marcos Ortega"
        ],
        "subjectives": [
          "Image and Video Processing (eess.IV)",
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      },
      {
        "id": "arXiv:2212.00920",
        "title": "Geometry-Aware Network for Domain Adaptive Semantic Segmentation",
        "abstract": " Comments: AAAI 2023 ",
        "url": "https://arxiv.org/abs/2212.00920",
        "authors": [
          "Yinghong Liao",
          "Wending Zhou",
          "Xu Yan",
          "Shuguang Cui",
          "Yizhou Yu",
          "Zhen Li"
        ],
        "subjectives": [
          "Computer Vision and Pattern Recognition (cs.CV)"
        ]
      }
    ]

    Comments

    More Rules

    View all

    Textbased Predictor Perplexity Rules

    H
    Havcker243

    Intelligent Team Building Recommendation System Perplexity Rules

    R
    royswastik

    Pita Perplexity Rules

    S
    SaratBobbili

    Eskrev Perplexity Rules

    R
    rfmss

    Founder Control Room Perplexity Rules

    J
    jussray

    N8nWorkflows Perplexity Rules

    S
    Samarth-ITM

    Stay up to date

    Get the latest Perplexity prompts, rules, and resources delivered to your inbox weekly.

    Neura Market LogoNeura Market

    Discover the best AI prompts, plugins, and resources for Perplexity and more.

    Content Types

    • Rules
    • Prompts
    • MCPs
    • Agents
    • Guides

    Platforms

    • ChatGPT Directory
    • Claude Directory
    • Gemini Directory
    • Cursor Directory
    • Grok Directory
    • Perplexity Directory
    • DeepSeek Directory
    • CoPilot Directory
    • Stable Diffusion Directory
    • Midjourney Directory
    • All Directories

    Resources

    • Blog
    • Documentation
    • Help Center
    • Marketplace

    Legal

    • Privacy Policy
    • Terms of Service

    © 2026 Neura Market. All rights reserved.

    |

    Not affiliated with any AI platform vendors.

    Neura Market

    Custom AI Systems & Services

    Our team of experienced AI builders will help build custom AI systems, workflows, and solutions for your business.

    Request custom work

    Ready-made automations for this

    Workflows from the Neura Market marketplace related to this Perplexity resource

    • Daily Auto-Generated Tweets from Trending Topics using Perplexity & GPT-4n8n · $4.99 · Related topic
    • Automated Lead Generation & Contact Enrichment with Hunter.io and Perplexity AIn8n · $24.99 · Related topic
    • Automate SEO-Optimized Blog Creation with GPT-4, Perplexity AI & Multi-Language Supportn8n · $24.99 · Related topic
    • Automate SEO Blog Content Creation with GPT-4, Perplexity AI, and WordPressn8n · $24.99 · Related topic
    Browse all workflows