nickmoranp.bsky.social
@nickmoranp.bsky.social
PrivateBin OSS pastebin where the server has zero knowledge of pasted data - https://privatebin.info/
PrivateBin
General information on the PrivateBin project.
privatebin.info
November 18, 2024 at 8:42 PM
Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement - https://huggingface.co/papers/2411.06558
Paper page - Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement
Join the discussion on this paper page
huggingface.co
November 18, 2024 at 5:27 PM
GaussianAnything: Interactive Point Cloud Latent Diffusion for 3D Generation - https://huggingface.co/papers/2411.08033
Paper page - GaussianAnything: Interactive Point Cloud Latent Diffusion for 3D Generation
Join the discussion on this paper page
huggingface.co
November 18, 2024 at 5:27 PM
The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use - https://huggingface.co/papers/2411.10323
Paper page - The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use
Join the discussion on this paper page
huggingface.co
November 18, 2024 at 5:27 PM
LLaVA-o1: Let Vision Language Models Reason Step-by-Step - https://huggingface.co/papers/2411.10440
Paper page - LLaVA-o1: Let Vision Language Models Reason Step-by-Step
Join the discussion on this paper page
huggingface.co
November 18, 2024 at 5:27 PM
Number it: Temporal Grounding Videos like Flipping Manga - https://huggingface.co/papers/2411.10332
Paper page - Number it: Temporal Grounding Videos like Flipping Manga
Join the discussion on this paper page
huggingface.co
November 18, 2024 at 5:27 PM
Direct Preference Optimization Using Sparse Feature-Level Constraints - https://huggingface.co/papers/2411.07618
Paper page - Direct Preference Optimization Using Sparse Feature-Level Constraints
Join the discussion on this paper page
huggingface.co
November 14, 2024 at 8:43 PM
Large Language Models Can Self-Improve in Long-context Reasoning - https://huggingface.co/papers/2411.08147
Paper page - Large Language Models Can Self-Improve in Long-context Reasoning
Join the discussion on this paper page
huggingface.co
November 14, 2024 at 8:43 PM
EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation - https://huggingface.co/papers/2411.08380
Paper page - EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation
Join the discussion on this paper page
huggingface.co
November 14, 2024 at 8:43 PM
CamemBERT 2.0: A Smarter French Language Model Aged to Perfection - https://huggingface.co/papers/2411.08868
Paper page - CamemBERT 2.0: A Smarter French Language Model Aged to Perfection
Join the discussion on this paper page
huggingface.co
November 14, 2024 at 8:43 PM
Can sparse autoencoders be used to decompose and interpret steering vectors? - https://huggingface.co/papers/2411.08790
Paper page - Can sparse autoencoders be used to decompose and interpret steering vectors?
Join the discussion on this paper page
huggingface.co
November 14, 2024 at 8:43 PM
RESOLVE: Relational Reasoning with Symbolic and Object-Level Features Using Vector Symbolic Processing - https://arxiv.org/abs/2411.08290
RESOLVE: Relational Reasoning with Symbolic and Object-Level Features Using Vector Symbolic Processing
Modern transformer-based encoder-decoder architectures struggle with reasoning tasks due to their inability to effectively extract relational information between input objects (data/tokens). Recent work introduced the Abstractor module, embedded between transformer layers, to address this gap. However, the Abstractor layer while excelling at capturing relational information (pure relational reasoning), faces challenges in tasks that require both object and relational-level reasoning (partial relational reasoning). To address this, we propose RESOLVE, a neuro-vector symbolic architecture that combines object-level features with relational representations in high-dimensional spaces, using fast and efficient operations such as bundling (summation) and binding (Hadamard product) allowing both object-level features and relational representations to coexist within the same structure without interfering with one another. RESOLVE is driven by a novel attention mechanism that operates in a bipolar high dimensional space, allowing fast attention score computation compared to the state-of-the-art. By leveraging this design, the model achieves both low compute latency and memory efficiency. RESOLVE also offers better generalizability while achieving higher accuracy in purely relational reasoning tasks such as sorting as well as partial relational reasoning tasks such as math problem-solving compared to state-of-the-art methods.
arxiv.org
November 14, 2024 at 7:54 PM
MVKTrans: Multi-View Knowledge Transfer for Robust Multiomics Classification - https://arxiv.org/abs/2411.08703
MVKTrans: Multi-View Knowledge Transfer for Robust Multiomics Classification
The distinct characteristics of multiomics data, including complex interactions within and across biological layers and disease heterogeneity (e.g., heterogeneity in etiology and clinical symptoms), drive us to develop novel designs to address unique challenges in multiomics prediction. In this paper, we propose the multi-view knowledge transfer learning (MVKTrans) framework, which transfers intra- and inter-omics knowledge in an adaptive manner by reviewing data heterogeneity and suppressing bias transfer, thereby enhancing classification performance. Specifically, we design a graph contrastive module that is trained on unlabeled data to effectively learn and transfer the underlying intra-omics patterns to the supervised task. This unsupervised pretraining promotes learning general and unbiased representations for each modality, regardless of the downstream tasks. In light of the varying discriminative capacities of modalities across different diseases and/or samples, we introduce an adaptive and bi-directional cross-omics distillation module. This module automatically identifies richer modalities and facilitates dynamic knowledge transfer from more informative to less informative omics, thereby enabling a more robust and generalized integration. Extensive experiments on four real biomedical datasets demonstrate the superior performance and robustness of MVKTrans compared to the state-of-the-art. Code and data are available at https://github.com/Yaolab-fantastic/MVKTrans.
arxiv.org
November 14, 2024 at 7:54 PM
A Fuzzy Reinforcement LSTM-based Long-term Prediction Model for Fault Conditions in Nuclear Power Plants - https://arxiv.org/abs/2411.08370
A Fuzzy Reinforcement LSTM-based Long-term Prediction Model for Fault Conditions in Nuclear Power Plants
Early fault detection and timely maintenance scheduling can significantly mitigate operational risks in NPPs and enhance the reliability of operator decision-making. Therefore, it is necessary to develop an efficient Prognostics and Health Management (PHM) multi-step prediction model for predicting of system health status and prompt execution of maintenance operations. In this study, we propose a novel predictive model that integrates reinforcement learning with Long Short-Term Memory (LSTM) neural networks and the Expert Fuzzy Evaluation Method. The model is validated using parameter data for 20 different breach sizes in the Main Steam Line Break (MSLB) accident condition of the CPR1000 pressurized water reactor simulation model and it demonstrates a remarkable capability in accurately forecasting NPP parameter changes up to 128 steps ahead (with a time interval of 10 seconds per step, i.e., 1280 seconds), thereby satisfying the temporal advance requirement for fault prognostics in NPPs. Furthermore, this method provides an effective reference solution for PHM applications such as anomaly detection and remaining useful life prediction.
arxiv.org
November 14, 2024 at 7:54 PM
AstroM$^3$: A self-supervised multimodal model for astronomy - https://arxiv.org/abs/2411.08842
AstroM$^3$: A self-supervised multimodal model for astronomy
While machine-learned models are now routinely employed to facilitate astronomical inquiry, model inputs tend to be limited to a primary data source (namely images or time series) and, in the more advanced approaches, some metadata. Yet with the growing use of wide-field, multiplexed observational resources, individual sources of interest often have a broad range of observational modes available. Here we construct an astronomical multimodal dataset and propose AstroM$^3$, a self-supervised pre-training approach that enables a model to learn from multiple modalities simultaneously. Specifically, we extend the CLIP (Contrastive Language-Image Pretraining) model to a trimodal setting, allowing the integration of time-series photometry data, spectra, and astrophysical metadata. In a fine-tuning supervised setting, our results demonstrate that CLIP pre-training improves classification performance for time-series photometry, where accuracy increases from 84.6% to 91.5%. Furthermore, CLIP boosts classification accuracy by up to 12.6% when the availability of labeled data is limited, showing the effectiveness of leveraging larger corpora of unlabeled data. In addition to fine-tuned classification, we can use the trained model in other downstream tasks that are not explicitly contemplated during the construction of the self-supervised model. In particular we show the efficacy of using the learned embeddings for misclassifications identification, similarity search, and anomaly detection. One surprising highlight is the "rediscovery" of Mira subtypes and two Rotational variable subclasses using manifold learning and dimension reduction algorithm. To our knowledge this is the first construction of an $n>2$ mode model in astronomy. Extensions to $n>3$ modes is naturally anticipated with this approach.
arxiv.org
November 14, 2024 at 7:54 PM
DNN Task Assignment in UAV Networks: A Generative AI Enhanced Multi-Agent Reinforcement Learning Approach - https://arxiv.org/abs/2411.08299
DNN Task Assignment in UAV Networks: A Generative AI Enhanced Multi-Agent Reinforcement Learning Approach
Unmanned Aerial Vehicles (UAVs) possess high mobility and flexible deployment capabilities, prompting the development of UAVs for various application scenarios within the Internet of Things (IoT). The unique capabilities of UAVs give rise to increasingly critical and complex tasks in uncertain and potentially harsh environments. The substantial amount of data generated from these applications necessitates processing and analysis through deep neural networks (DNNs). However, UAVs encounter challenges due to their limited computing resources when managing DNN models. This paper presents a joint approach that combines multiple-agent reinforcement learning (MARL) and generative diffusion models (GDM) for assigning DNN tasks to a UAV swarm, aimed at reducing latency from task capture to result output. To address these challenges, we first consider the task size of the target area to be inspected and the shortest flying path as optimization constraints, employing a greedy algorithm to resolve the subproblem with a focus on minimizing the UAV's flying path and the overall system cost. In the second stage, we introduce a novel DNN task assignment algorithm, termed GDM-MADDPG, which utilizes the reverse denoising process of GDM to replace the actor network in multi-agent deep deterministic policy gradient (MADDPG). This approach generates specific DNN task assignment actions based on agents' observations in a dynamic environment. Simulation results indicate that our algorithm performs favorably compared to benchmarks in terms of path planning, Age of Information (AoI), energy consumption, and task load balancing.
arxiv.org
November 14, 2024 at 7:54 PM
Towards Optimizing a Retrieval Augmented Generation using Large Language Model on Academic Data - https://arxiv.org/abs/2411.08438
Towards Optimizing a Retrieval Augmented Generation using Large Language Model on Academic Data
Given the growing trend of many organizations integrating Retrieval Augmented Generation (RAG) into their operations, we assess RAG on domain-specific data and test state-of-the-art models across various optimization techniques. We incorporate four optimizations; Multi-Query, Child-Parent-Retriever, Ensemble Retriever, and In-Context-Learning, to enhance the functionality and performance in the academic domain. We focus on data retrieval, specifically targeting various study programs at a large technical university. We additionally introduce a novel evaluation approach, the RAG Confusion Matrix designed to assess the effectiveness of various configurations within the RAG framework. By exploring the integration of both open-source (e.g., Llama2, Mistral) and closed-source (GPT-3.5 and GPT-4) Large Language Models, we offer valuable insights into the application and optimization of RAG frameworks in domain-specific contexts. Our experiments show a significant performance increase when including multi-query in the retrieval phase.
arxiv.org
November 14, 2024 at 7:54 PM
RLInspect: An Interactive Visual Approach to Assess Reinforcement Learning Algorithm - https://arxiv.org/abs/2411.08392
RLInspect: An Interactive Visual Approach to Assess Reinforcement Learning Algorithm
Reinforcement Learning (RL) is a rapidly growing area of machine learning that finds its application in a broad range of domains, from finance and healthcare to robotics and gaming. Compared to other machine learning techniques, RL agents learn from their own experiences using trial and error, and improve their performance over time. However, assessing RL models can be challenging, which makes it difficult to interpret their behaviour. While reward is a widely used metric to evaluate RL models, it may not always provide an accurate measure of training performance. In some cases, the reward may seem increasing while the model's performance is actually decreasing, leading to misleading conclusions about the effectiveness of the training. To overcome this limitation, we have developed RLInspect - an interactive visual analytic tool, that takes into account different components of the RL model - state, action, agent architecture and reward, and provides a more comprehensive view of the RL training. By using RLInspect, users can gain insights into the model's behaviour, identify issues during training, and potentially correct them effectively, leading to a more robust and reliable RL system.
arxiv.org
November 14, 2024 at 7:54 PM
PerceiverS: A Multi-Scale Perceiver with Effective Segmentation for Long-Term Expressive Symbolic Music Generation - https://arxiv.org/abs/2411.08307
PerceiverS: A Multi-Scale Perceiver with Effective Segmentation for Long-Term Expressive Symbolic Music Generation
Music generation has progressed significantly, especially in the domain of audio generation. However, generating symbolic music that is both long-structured and expressive remains a significant challenge. In this paper, we propose PerceiverS (Segmentation and Scale), a novel architecture designed to address this issue by leveraging both Effective Segmentation and Multi-Scale attention mechanisms. Our approach enhances symbolic music generation by simultaneously learning long-term structural dependencies and short-term expressive details. By combining cross-attention and self-attention in a Multi-Scale setting, PerceiverS captures long-range musical structure while preserving performance nuances. The proposed model, evaluated on datasets like Maestro, demonstrates improvements in generating coherent and diverse music with both structural consistency and expressive variation. The project demos and the generated music samples can be accessed through the link: https://perceivers.github.io.
arxiv.org
November 14, 2024 at 7:54 PM