Journal articleACM Transactions on Architecture and Code Optimization · September 17, 2025
Markov-Chain Monte-Carlo (MCMC) algorithms offer a general framework for performing interpretable inference but have high overheads due to the computational complexity of the sampling process and the large number of samples required to produce an accurate ...
Full textCite
Journal articleACM Transactions on Graphics · December 19, 2024
The growing demands for computational power in cloud computing have led to a significant increase in the deployment of high-performance servers. The growing power consumption of servers and the heat they produce is on track to outpace the capacity of conve ...
Full textCite
Journal articleIEEE Micro · January 1, 2023
Remote direct memory access (RDMA) networks enable low latency and low central processing unit utilization, and their widespread adoption in datacenters enables improved application performance. However, there are performance isolation concerns for RDMA de ...
Full textCite
Journal articleNano letters · June 2017
We demonstrate an optically controlled molecular-scale pass gate that uses the photoinduced dark states of fluorescent molecules to modulate the flow of excitons. The device consists of four fluorophores spatially arranged on a self-assembled DNA nanostruc ...
Full textCite
Journal articleIEEE Micro · January 1, 2017
As lithographic feature sizes approach fundamental scaling limits, a variety of computational domains remain incompatible with integrated circuits merely due to their operating principles. Resonance energy transfer (RET) logic offers a molecular-scale solu ...
Full textCite
Journal articleIEEE Micro · September 1, 2015
Despite the theoretical advances in probabilistic computing, a fundamental mismatch persists between the deterministic hardware that traditional computers use and the stochastic nature of probabilistic algorithms. In this article, the authors propose Reson ...
Full textOpen AccessCite
Journal articleInternational Conference on Architectural Support for Programming Languages and Operating Systems ASPLOS · March 14, 2015
Molecular-scale Network-on-Chip (mNoC) crossbars use quantum dot LEDs as an on-chip light source, and chromophores to provide optical signal filtering for receivers. An mNoC reduces power consumption or enables scaling to larger crossbars for a reduced ene ...
Full textOpen AccessCite
Journal articleJournal of Parallel and Distributed Computing · January 1, 2014
Optical nanoscale computing is one promising alternative to the CMOS process. In this paper we explore the application of Resonance Energy Transfer (RET) logic to common digital circuits. We propose an Optical Logic Element (OLE) as a basic unit from which ...
Full textCite
Journal articleIEEE Micro · January 1, 2011
Computer systems with virtual memory are susceptible to design bugs and runtime faults in their address translation systems. Detecting bugs and faults requires a clear specification of correct behavior. A new framework for address translation aware memory ...
Full textCite
Journal articleIEEE Computer Architecture Letters · July 1, 2010
One of the most challenging problems in developing a multicore processor is verfiying that the design is correct, and one of the most difficult aspects of pre-silicon verification is verifying that the memory system obeys the architecture’s specified ...
Full textOpen AccessCite
Journal articleSmall (Weinheim an der Bergstrasse, Germany) · April 2010
The self-assembly of molecularly precise nanostructures is widely expected to form the basis of future high-speed integrated circuits, but the technologies suitable for such circuits are not well understood. In this work, DNA self-assembly is used to creat ...
Full textCite
Journal articleACM Journal on Emerging Technologies in Computing Systems · March 1, 2010
The integration of novel nanotechnologies onto silicon platforms is likely to increase fabrication defects compared with traditional CMOS technologies. Furthermore, the number of nodes connected with these networks makes acquiring a global defect map impra ...
Full textCite
Journal articleIEEE Micro · January 1, 2010
The authors explore nanoscale sensor processor (nSP) architectures. Their design includes a simple accumulator-based instruction-set architecture, sensors, limited memory, and instruction-fused sensing. Using nSP technology based on optical resonance energ ...
Full textOpen AccessCite
Journal articleInternational Conference on Architectural Support for Programming Languages and Operating Systems ASPLOS · March 7, 2009
This paper explores the architectural implications of integrating computation and molecular probes to form nanoscale sensor processors (nSP). We show how nSPs may enable new computing domains and automate tasks that currently require expert scientific trai ...
Full textCite
Journal articleIEEE Micro · January 1, 2008
Drawing on the nanometer-placement capabilities of self-assembly fabrication methods, the authors propose a new nanoscale device based on a single-molecule optical phenomenon called resonance energy transfer. This device enables a complete integrated techn ...
Full textCite
Journal articleACM Journal on Emerging Technologies in Computing Systems · July 1, 2007
The continual decrease in transistor size (through either scaled CMOS or emerging nanotechnologies) promises to usher in an era of tera to peta-scale integration but with increasing defects. Regardless of fabrication methodology (top-down or bottom-up), de ...
Full textCite
Journal articleInternational Conference on Architectural Support for Programming Languages and Operating Systems ASPLOS · October 23, 2006
The continual decrease in transistor size (through either scaled CMOS or emerging nano-technologies) promises to usher in an era of tera to peta-scale integration. However, this decrease in size is also likely to increase defect densities, contributing to ...
Full textCite
Journal articleIEEE Transactions on Parallel and Distributed Systems · June 1, 2006
Spinning is a synchronization mechanism commonly used in applications and operating systems. Excessive spinning, however, often indicates performance or correctness (e.g., livelock) problems. Detecting if applications and operating systems are spinning is ...
Full textCite
Journal articleACM Journal on Emerging Technologies in Computing Systems · January 1, 2006
This article explores the architectural challenges introduced by emerging bottom-up fabrication of nanoelectronic circuits. The specific nanotechnology we explore proposes patterned DNA nanostructures as a scaffold for the placement and interconnection of ...
Full textCite
Journal articleProceedings Design Automation Conference · January 1, 2006
DNA self-assembly is an emerging technology with potential as a future replacement of conventional lithographic fabrication. A key challenge is the specification of appropriate DNA sequences that are optimal according to specified metrics and satisfy vario ...
Full textCite
Journal articleComputer · January 1, 2005
Despite the convenience of clean abstractions, technological trends are blurring the lines between design layers and creating new interactions between previously unrelated architecture layers. For example, virtual machines such as VMWare and Transmeta impl ...
Full textCite
Journal articleNanotechnology · September 1, 2004
The shift in technology away from silicon complementary metal-oxide semiconductors (CMOS) to novel nanoscale technologies requires new design tools. In this paper, we explore one particular nanotechnology: carbon nanotube transistors that are self-assemble ...
Full textCite
Journal articleIEEE Trans. Parallel Distrib. Syst. (USA) · 2004
Network performance in tightly-coupled multiprocessors typically degrades rapidly beyond network saturation. Consequently, designers must keep a network below its saturation point by reducing the load on the network. Congestion control via source throttlin ...
Full textLink to itemCite
Journal articleLecture Notes in Computer Science Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics · January 1, 2004
Energy consumption is becoming a limiting factor in the development of computer systems for a range of application domains. Since processor performance comes with a high power cost, there is increased interest in scaling the CPU voltage and clock frequency ...
Cite
Journal articleACM Transactions on Architecture and Code Optimization · January 1, 2004
Prefetching is often used to overlap memory latency with computation for array-based applications. However, prefetching for pointer-intensive applications remains a challenge because of the irregular memory access pattern and pointer-chasing problem. In th ...
Full textCite
Journal articleIEEE Transactions on Parallel and Distributed Systems · November 1, 2002
The performance of both serial and parallel implementations of matrix multiplication is highly sensitive to memory system behavior. False sharing and cache conflicts cause traditional column-major or row-major array layouts to incur high variability in mem ...
Full textCite
Journal articleJournal of Computer and System Sciences · January 1, 2001
In this paper we construct an analytic model of cache misses during matrix multiplication. The analysis in this paper applies to square matrices of size m where the array layout function is given in terms of a function Θ that interleaves the bit ...
Full textCite
Journal articleIEEE Trans. Comput. (USA) · September 2000
Three-dimensional (3D) graphics applications have become very important workloads running on today's computer systems. A cost-effective graphics solution is to perform geometry processing of 3D graphics on the host CPU and have specialized hardware handle ...
Full textLink to itemCite
Journal articleJournal of Instruction Level Parallelism · October 1, 1999
This paper provides a quantitative evaluation of load latency tolerance in a dynamically scheduled processor. To determine the latency tolerance of each memory load operation, our simulations use flexible load completion policies instead of a fixed memory ...
Cite
Journal articleACM Transactions on Modeling and Computer Simulation · January 1, 1997
This article describes the active memory abstraction for memory-system simulation. In this abstraction - designed specifically for on-the-fly simulation - memory references logically invoke a user-specified function depending upon the reference's type and ...
Full textCite
Journal articleIEEE Transactions on Parallel and Distributed Systems · January 1, 1994
Several techniques have been proposed to allow parallel access to a shard memory location by combining requests. They have one or more of the following attributes: requirements for a priori knowledge of the request to combine, restrictions on the routing o ...
Full textCite
Journal articleProceedings of the 1993 ACM Sigmetrics Conference on Measurement and Modeling of Computer Systems Sigmetrics 1993 · June 1, 1993
We have developed a new technique for evaluating cache coherent, shared-memory computers. The Wisconsin Wind Tunnel (WWT) runs a parallel shared-memory program on a parallel computer (CM-5) and uses execution-driven, distributed, discrete-event simulation ...
Full textCite
Journal articleComput. Archit. News (USA) · January 9, 1993
Wisconsin Architectural Research Tool Set (WARTS) is a collection of tools for profiling and tracing programs and analyzing program traces. WARTS currently contains: QPT, a program profiler and tracing system; CPROF, a cache performance profiler; and Tycho ...
Full textLink to itemCite