Tags¶
Following is a list of relevant tags:
0/1 Adam¶
1-bit Adam¶
1-bit LAMB¶
5G¶
64bits vs 32bits¶
AGRS¶
AI¶
- AI Compiler
- AI Image
- Agile Governance: Balancing IPD and AI Innovation
- Code Migration And Alignment
- IPD Q&A
- Introduction to AI and Machine Learning Basics
- Subagent Control Plane
- 成功的软件商业解决方案
AI AGENT¶
AI Cluster¶
AI INFRA¶
AI Infrastructure¶
AI STARTUP¶
AIInfra¶
AISBench¶
ALST¶
AMAT¶
AMD¶
APR¶
ASPLOS¶
ATI¶
ATTENTION RESIDUALS¶
AVX¶
Activation Offload¶
Agent¶
Agentic RL¶
Algorithm¶
All-to-All¶
AllGather¶
Alpha¶
Analysis¶
Apt¶
Archon¶
Ascend¶
- BSND TND Operator Layout
- CloudMatrix384 LLM Serving
- ClusterHealthDetect A3 Performance
- Memory-Semantic TransferQueue
- NPU Training Operators - GDN
- NPU Training Operators - GMM
- NPU Training Operators - MC2
- NPU Training Operators - RoPE MRoPE
- Next of My Ascend Career
Ascend 910C¶
Ascend 950¶
Ascend A5¶
Assembly¶
Async¶
Async RL¶
Attention¶
AudioVideo¶
AutoResearch¶
AutoTP¶
Autonomous Driving¶
Autotuning¶
BFS¶
BHive¶
BSHD¶
BSND¶
BT¶
BTL¶
Baka Mitai¶
Bash¶
Big-Endian¶
Blackwell¶
C¶
- Naming
- [C++ Basic] Exploring Useful Built-in Functions
- [C++ Basic] Grammar
- [C++ Basic] Types
- [C++] Destructor Order
C++¶
- Gperftools
- Hash map
- Naming
- [C++ Basic] Exploring Useful Built-in Functions
- [C++ Basic] Grammar
- [C++ Basic] STL Data Structure
- [C++ Basic] Types
- [C++ Basic] User-Defined Types: Class
- [C++] Destructor Order
CCD¶
CCX¶
CI¶
CISC¶
COMPENSATION¶
CP¶
CPP¶
CPU¶
CPU Thread Socket¶
CUDA¶
CUDA Allocator¶
CV¶
Calling Conventions¶
Career¶
- 1.2 Career
- 1.2 Career:1 秋招
- Career Transferable skill / Durable skills / Core capabilities
- DFX: Design for X
- Multi-Objective Decision Making
- QCC:Quality Control Circle
Career Strategy¶
Checkpoint¶
Chunk Layer¶
Class¶
CloudMatrix384¶
Cloudflare¶
ClusterHealthDetect¶
Codeforces¶
Collective Communication¶
Communication¶
Communication Logging¶
Communication-Compute Fusion¶
Compiler¶
Conference¶
Context Parallelism¶
Crawler¶
DDR¶
DFX¶
DFlash¶
DLP¶
DNS¶
DP¶
DPO¶
DRAM¶
DSA¶
DanceGRPO¶
DanceOPD¶
DataStates¶
Database¶
Databases¶
Dataset¶
Debug¶
Decision Analysis¶
Deep Learning¶
DeepEP¶
DeepNVMe¶
DeepSeek-R1¶
DeepSpeed¶
- DeepSpeed Communication Compression and Hiding
- DeepSpeed I/O, Offload, and Asynchrony
- DeepSpeed Memory and Parallelism
- DeepSpeed MoE and Model Compression
- DeepSpeed Observability and Autotuning
Deepfake¶
DevLog¶
- [DevLog] PLAN
- [DevLog]24Q3P1 - Optimize PTA With Thread Affinity
- [DevLog]24Q4P2 - Lazy Initialize during old dispatcher way
DiT¶
Diffusion¶
Diffusion Language Model¶
Diffusion Model¶
Digital Worker¶
- AI Documentation Workflow
- My Digital Worker
- My Digital Worker : AutoMoneyMaker - AutoTrader
- My Digital Worker : Model / Software Usage
- My Digital Worker : New Coding Way
- My Digital Worker : New Coding Way Part0 —— Building AI-Coding Env
- My Digital Worker : Work with AI
- Personal Advantage Workflow
Disk¶
Domino¶
Domino EP¶
Dynamic Programming¶
ECC¶
ESOP¶
Echart¶
Ed2k¶
Evaluation¶
Executable file¶
Expert Parallelism¶
- CloudMatrix384 LLM Serving
- NPU Training Operators - GMM
- NPU Training Operators - MC2
- XTuner Domino EP
Explain¶
FLA¶
FLOPs Profiler¶
FMA¶
FP8¶
FPDT¶
FSDP¶
FSDP2¶
FeatureMatrix¶
FlashInfer¶
Flow-Factory¶
FlowServe¶
Foundation Model Training¶
Frontier Models¶
FurtherStudy¶
GCN¶
GDDR6x¶
GDN¶
- Attention Architecture Evolution
- Attention Cache and Sequence Parallelism
- BSND TND Operator Layout
- NPU Training Operators - GDN
- Training Performance Model
- vLLM Inference Profiling
GIL¶
GLIBC¶
GLM-5.2¶
GMM¶
GNN¶
GNU¶
GPT¶
GPTQ¶
GPU¶
GQA¶
GRPO¶
- BSND TND Operator Layout
- Diffusion LLM Post-Training
- NPU Training Operators - GDN
- RL Algorithms: PPO-RLHF & GRPO-family
Game¶
H2D¶
HCCL¶
HDMI¶
HPCAI¶
HPL¶
HPL-PL¶
Hopper¶
Huawei¶
Hugo¶
HybridEP¶
ILP¶
IP¶
IPCC¶
IPD¶
IPO¶
Image¶
Inference¶
Inference Quantization¶
Jellyfin¶
Jupyter¶
KDA¶
KIMI¶
KV Cache¶
Kavita¶
Kimi K3¶
Knowledge Distillation¶
Komga¶
Kunpeng¶
LATENTMOE¶
LDA¶
LLM¶
LLM Inference¶
LLM Serving¶
LLM Wiki¶
LLaDA¶
LeetCode¶
Leetcode¶
LegacyBugs¶
Linking¶
Lock¶
MAC¶
MC2¶
MCA¶
MCDA¶
MFU¶
MHA¶
MILAN¶
MIPS¶
MLA¶
MLP¶
MOE¶
MOPD¶
MPI¶
- Dynamic pool dispatch
- IPCC Preliminary SLIC Optimization 5: MPI + OpenMP
- IPCC Preliminary SLIC Optimization 6: Non-blocking MPI
- MPI
- Memory Semantics vs RDMA
- Python MPI
- Why MPI_Init is slow
MPI option¶
MPI_Init¶
MPK¶
MPS¶
MQA¶
MRoPE¶
MTE¶
MTP¶
MULTIMODAL¶
MXFP8¶
Mac¶
Machine Learning¶
MegaKernel¶
MegaMoE¶
Megatron¶
Megatron Bridge¶
Megatron Core¶
MemFabric¶
Memory Optimization¶
Memory Semantics¶
Mesh Interconnect Architecture¶
Metrics¶
Micro¶
Micro-Fusion¶
Micro-architecture¶
Microarchitecture¶
- Microarchitecture: Micro-Fusion & Macro-Fusion
- Microarchitecture: Out-Of-Order execution(OoOE/OOE) & Register Renaming
- Microarchitecture: Pipeline of Intel Core CPUs
- Microarchitecture: Zero (one) idioms & Mov Elimination
MindSpeed¶
MindSpeed-MM¶
- AI Training Parallelism
- NPU Training Operators - GMM
- NPU Training Operators - MC2
- VeRL Backend Parallelism
MixZ++¶
MoE¶
- CloudMatrix384 LLM Serving
- DeepSpeed MoE and Model Compression
- Inference MegaKernel
- Kimi K3 NPU Training
- NPU Training Operators - GMM
- NPU Training Operators - MC2
- Training Performance Model
- VeRL Router Replay
- XTuner Domino EP
- XTuner Memory Optimization
- xDeepServe on CloudMatrix384
MoQ¶
Model Compression¶
Module¶
Money¶
Monitor¶
Mount¶
Mov-Elimination¶
Multi-Agent¶
Multimodal¶
Multimodal Generation¶
Multimodal Model¶
Muon¶
NAS¶
NFT¶
NLP¶
NPU¶
- AI Infra Daily Radar
- BSND TND Operator Layout
- Kimi K3 NPU Training
- NPU Training Operators - GDN
- NPU Training Operators - GMM
- NPU Training Operators - MC2
- NPU Training Operators - RoPE MRoPE
- VeRL Async Policy
- VeRL Performance Optimization
NUMA¶
NV¶
NVLink¶
Nas¶
Nvidia¶
O2¶
O3¶
OPD¶
OS¶
OSS¶
Omni¶
OmniNFT¶
Open Source¶
OpenLDAP¶
OpenMP¶
OpenWRT¶
Openssl¶
Options¶
- AMD Epyc Compiler Options
- GCC Compiler Option 1 : Optimization Options
- GCC Compiler Option 2 : Preprocessor Options
- Intel Compile Options
OutOfOrder¶
PCIe¶
PIM¶
PKI¶
PML¶
PP¶
PPO¶
PPT¶
PT¶
PTA¶
- Debug/Profile/Devlop Tools of PTA
- [DevLog] PLAN
- [DevLog]24Q3P1 - Optimize PTA With Thread Affinity
- [DevLog]24Q4P2 - Lazy Initialize during old dispatcher way
PTX¶
PVE¶
Pagerank¶
Parallel¶
Parallelism¶
Perf¶
Performance¶
Performance Engineering¶
Performance Model¶
Performance Modeling¶
Pi¶
PicBed¶
PicGo¶
Pin¶
Pipeline¶
Post Training¶
Post-Training¶
PostTraining¶
Powershell¶
Presentation¶
Priority¶
Probability Theory¶
Professional Skills¶
PyG¶
PyPTO¶
PyTorch¶
PyTorch Profiler¶
Pytorch¶
- Pytorch 1 :Basic Components
- Pytorch 2 :more conceptions about training and inference
- Pytorch 2.5 :Dataset & Dataloader
- Pytorch 3 :Model & Training
- Pytorch 4 :Save & Load & Pretrain
- Pytorch 5 : Distributed Training & Parallelism (Mem Re-allocated)
- Pytorch 6 :Visualization
- Pytorch 7 :Memory Optimization(Freeing GPU/NPU Memory Early)
- Pytorch 8 :Hyperparameter
QCC¶
QuantitativeFinance¶
Quantization¶
Qwen3-Omni¶
Qwen3-VL¶
Qwen3.5¶
- BSND TND Operator Layout
- Inference Quantization Formats
- NPU Training Operators - GDN
- NPU Training Operators - GMM
- Training Performance Model
RAG¶
RAID¶
RAW¶
RDMA¶
RGB Lab¶
RHLF¶
RISC¶
RISC-V¶
RL¶
- AI Post Traning: DPO + MPO
- AI Post Traning: DanceGRPO
- Agent & Agentic RL
- DiffusionNFT
- Fast Debug: VeRL example
- Frontier Model RL
- Kimi K3 Report
- Memory-Semantic TransferQueue
- Multimodal Generation Evaluation
- Multimodal RL
- RL Algorithms: PPO-RLHF & GRPO-family
- RL DFX Metrics
- RL Data Flow
- RL Infra Series
- RL Next: Meta-Learning
- RL: Training Inference Mismatch
- RL: xPU Mismatch - metrics
- The Mechanics of RL: How Inference Sampling Shapes the Probability Landscape
- Train Stages: Pretrain, Mid-Train(CT), SFT, RL
- VLM RL Evaluation Datasets
- VeRL
- VeRL Async
- VeRL Backend Parallelism
- VeRL Checkpoint
- VeRL Feature Matrix
- VeRL Feature Survey
- VeRL Performance Optimization
- VeRL Rollout Inference
- VeRL Router Replay
- VeRL Speculative Decoding
- VeRL Training Flow
RLHF¶
RLInfra¶
ROMA¶
RTX 3090¶
Ramulator¶
Register-Renaming¶
Reinforcement Learning¶
Reporting¶
Ring Attention¶
Risk Management¶
RoPE¶
RouterReplay¶
Rust¶
SASS¶
SE¶
SFT¶
SHMEM¶
SIMD¶
SLIC¶
- Hybrid Multithreaded/OpenMP + MPI parallel Programs
- IPCC Preliminary SLIC Analysis
- IPCC Preliminary SLIC Analysis part2 : Run process
- IPCC Preliminary SLIC Analysis part3 : Hot spot analysis
- IPCC Preliminary SLIC Analysis part4 : cluster environment
- IPCC Preliminary SLIC Case1/2/3
- IPCC Preliminary SLIC Optimization 2
- IPCC Preliminary SLIC Optimization 3
- IPCC Preliminary SLIC Optimization 4: EnforceLabelConnectivity
- IPCC Preliminary SLIC Optimization 5: MPI + OpenMP
- IPCC Preliminary SLIC Optimization 6: Non-blocking MPI
- IPCC Preliminary SLIC algorithm
- IPCC Preliminary SLIC test
SMA¶
SP¶
SSE¶
STEPFUN¶
SWAP¶
Scaling Law¶
Server¶
Services¶
Skylake¶
Slurm¶
SpeculativeDecoding¶
Staleness¶
Stanford¶
SuperPixel¶
Survey¶
Systemclt¶
Systemd¶
TLB¶
TLP¶
TND¶
TP¶
Terminal¶
Thunderbolt¶
Topcoder¶
Training¶
TransferQueue¶
Transformerless¶
Triton¶
Type-C¶
UB-Mesh¶
UPI¶
USP¶
Ubuntu¶
Ulysses¶
Ulysses-Offload¶
UniGRPO¶
UnifiedBus¶
Unreal¶
Useful¶
VBench¶
VEFX-Bench¶
VLA¶
VLM¶
- Frontier Model RL
- Ideas around Vision-Language Models (VLMs) / Reasoning Models
- Multimodal Generation Evaluation
- VLA VLM + DiT
- VLM RL Evaluation Datasets
VNC¶
VP¶
VPN¶
VR¶
Value Capture¶
Value of Information¶
VeOmni¶
VeRL¶
- AI Infra Daily Radar
- BSND TND Operator Layout
- Memory-Semantic TransferQueue
- NPU Training Operators - GDN
- Training Performance Model
- VLM RL Evaluation Datasets
- VeRL Async
- VeRL Async Policy
- VeRL Backend Parallelism
- VeRL Checkpoint
- VeRL Feature Matrix
- VeRL Feature Survey
- VeRL Performance Optimization
- VeRL Rollout Inference
- VeRL Router Replay
- VeRL Speculative Decoding
- VeRL Training Flow
VeRL-Omni¶
Vectorization¶
Video¶
Virtual memory¶
Visualization¶
Vscode¶
Vue¶
W8A4¶
W8A8¶
WAR¶
WAW¶
WLAN¶
Work Management¶
X86¶
XCCL¶
XTuner¶
Xeon¶
ZeRO++¶
ZeRO-Offload¶
ZenFlow¶
Zero-Idiom¶
Zsim¶
abi¶
agent¶
- My Digital Worker
- My Digital Worker : AutoMoneyMaker - AutoTrader
- My Digital Worker : Model / Software Usage
- My Digital Worker : New Coding Way
- My Digital Worker : New Coding Way Part0 —— Building AI-Coding Env
- My Digital Worker : Work with AI
ai¶
- AI Documentation Workflow
- AI Model Visualization
- Building Large-Scale AI Systems on Ascend: Training, Inference, and Multimodal Optimization
- Business Trip: 2601-2602 verl + DanceGRPO
- Deploy OpenLLM to one A100
- HPCAI
- My Digital Worker
- My Digital Worker : AutoMoneyMaker - AutoTrader
- My Digital Worker : Model / Software Usage
- My Digital Worker : New Coding Way
- My Digital Worker : New Coding Way Part0 —— Building AI-Coding Env
- My Digital Worker : Work with AI
- Personal Advantage Workflow
- Probability Theory
- Pytorch 1 :Basic Components
- Pytorch 2 :more conceptions about training and inference
- Pytorch 2.5 :Dataset & Dataloader
- Pytorch 3 :Model & Training
- Pytorch 4 :Save & Load & Pretrain
- Pytorch 5 : Distributed Training & Parallelism (Mem Re-allocated)
- Pytorch 6 :Visualization
- Pytorch 7 :Memory Optimization(Freeing GPU/NPU Memory Early)
- Pytorch 8 :Hyperparameter
algorithm¶
amd¶
ampere¶
anaconda¶
anime¶
- Anime Super Resolution to 4K & Interpolation to 120 fps
- Diary 230827: 上海二次元之旅
- UnimportantView: Anime Recommendation
- UnimportantView: Film & TV(Anime) Works Rating
aocc¶
apache¶
app¶
apple¶
apt¶
architecture¶
- GPU
- Microarchitecture: Micro-Fusion & Macro-Fusion
- Microarchitecture: Out-Of-Order execution(OoOE/OOE) & Register Renaming
- Microarchitecture: Pipeline of Intel Core CPUs
arm¶
array¶
ascend¶
assembly¶
auto¶
avx¶
avx256¶
bank¶
batch¶
benchmark¶
bilibili¶
blog¶
broadcast¶
bug¶
bugs¶
business¶
c¶
c++¶
cache¶
calloc¶
cat¶
centos¶
chatgpt¶
chivier¶
chsh¶
clang¶
clash¶
- Clash Config 4 yourself
- Clash on LAN/linux/Dockers
- ClashX Pro and Wireguard on Macbook In School Net
- OpenWRT on router
- Wireguard
class¶
cloud¶
clustering¶
cmake¶
code¶
color¶
comic¶
command¶
commands¶
compile¶
compile options¶
compression¶
conda¶
context switch¶
cpp¶
cpu¶
cpu flags¶
crontab¶
cs¶
css¶
ctags¶
cuda¶
- Cuda Optimize
- Cuda Optimize : Stencil
- Cuda Optimize : Vectorized Memory Access
- Cuda Program Basic
- Nvidia Nsight
- Nvprof
- The CUDA Execution Model
cuobjdump¶
dLLM¶
ddns¶
debug¶
deepseek¶
diagram¶
disassembly¶
disk¶
- Disk
- Disk C: make room for installation
- Migrate From Synology DS220J to UGREEN DX4600
- Mount Network Disk
- Nas Disk Speed Test
- Ubuntu server reInstall
dit¶
div¶
dnat¶
dnf¶
docker¶
docuwiki¶
dokuwiki¶
domain¶
domestic¶
dram¶
echart¶
email¶
english¶
enter¶
entertainment¶
epic¶
epyc¶
ethernet¶
family¶
filesystem¶
firewall¶
firewalld¶
flags¶
flood fill¶
focusk¶
fork¶
forward¶
fpic¶
fstab¶
ftp¶
fun¶
- (Research) Team/Lab Organization
- 0 Overview
- 1.1 Living Needs & Meaning
- 1: Target2chase
- 2.2 Social Part
- 2: Courage to move on
- 3 EfficientJumpingRunning
- 3.2 taskPriority
- 3.3 EfficientWorkLearning
- 6 FPS
- AI Hardware & Accelerators
- AI Infra: 10k-GPU cluster
- AI Training Optimization
- Anime Auto add Chinese Subtitle
- AntiCheat
- Audio
- Balance (Efficient) work & life time
- Benchmark
- Blind Date 1st
- Blind Date 1st(2)
- Blind Date Tips
- Blog writing
- Burnout Monitor : Healthy Body Model + hair/heart-aware exercise
- CProgramReading
- CV Model
- Chrome://tracing
- Classical AI Models
- Colorful Life (TOP)
- Cuda Driver Runtime
- Data Link
- Data Structure Summary
- Datastruture: Tree
- DeviceExpansion
- Disease And Prevention
- Disordered Ideas
- Excel
- Experiments For PIM Motivation
- FPGA
- Financial Values
- Future Plans: House & Car
- GUIAgents
- Go Templates
- HTML
- Host-Core With PIM-Core In 3D-stacked Mem
- Huawei Ascend Domain-Specific Architectures : DaVinci
- Important Date
- Inference Basic
- Inference Optimization
- Japanese
- Keyboard
- LLM Model
- LLM Model Basic
- Lab homepage Template & Website builders choice
- LinuxFolderInstall
- Mathematical Logic & Algebraic structure
- Mkdocs
- Muon Optimizer
- Muon Optimizer + FSDP
- OOTD: outfit of the day
- Open &Free Multimodel AI Tools
- OpenCL Basic
- OpenWRTNetworkManage
- Parallel_sort
- Personal Image Management
- Php
- Piano: Transcribe Piano Sheet Music from Video using AI model
- Postgraduate dormitory
- Predictor
- RL Weekly News
- Research logic
- Salary & Tax & Insurance
- Scientifically Concocting Glasses
- Search, Ads, and Recommendation
- Security
- Social Science
- Synology terminal
- TODO
- Team Cooperation / Relationship
- Tmux
- Topology
- Training Data Usage
- Turing Machine & P versus NP problem
- URLs
- UnimportantView: Film & TV(Anime) Works Rating
- UnimportantView: Game
- User Kernel Mode
- Wake On Lan(Wol)
- Weekly
- When & How 4 team presentation page & knowledge database pool
- 宛如泥潭的大型项目开发困境
function call¶
game¶
gateway¶
gcc¶
- C program compile&run process
- GCC Compile Error
- GCC Compiler Option 1 : Optimization Options
- GCC Compiler Option 2 : Preprocessor Options
- Inline Assembly
gdb¶
gdbgui¶
gef¶
gem5¶
gif¶
git¶
- Git Lfs
- Git Push 2 Homepage
- Git Standardization
- Git Submodule: Data & Code Repository Separate
- Introduction to Git Commands
- Network Basic
github¶
github action¶
glibc¶
gnu¶
go¶
- Go Install and Command
- Go mod
- Golang Syntax
- Web Design 2 : Content Organization & Link Content using go template
golang¶
gprof¶
gpt¶
gpu¶
graph¶
group¶
gzip¶
h265¶
hash¶
header¶
health¶
heap¶
hexo¶
home¶
homepage¶
- Git Push 2 Homepage
- Homepage Template Conflict
- How SSG Get Work? & Hugo theme creation
- Miscellaneous
- Web Server: Nginx V.S. Apache2
hpc¶
html¶
htop¶
http¶
huawei¶
- Forecasting Housing Prices Around Huawei Xicen Base: A Home Buying Plan for Qingpu District
- Huawei Kunpeng workload
- Travel and Business Trip Checklist
hugo¶
hw¶
icc¶
icecream¶
icpc¶
incomplete¶
- AOCC
- Cache
- GCC Compiler Option 1 : Optimization Options
- GCC Compiler Option 2 : Preprocessor Options
- Memalloc
inference¶
- SGLang
- Speculative Decoding & eagle3
- Vllm Basic
- vLLM Inference Profiling
- vllm-omni & DiT Inference Accelerate
initd¶
intel¶
- Intel Compile Options
- Intel Pin
- Intel SDM(Software Developer's Manual)
- Microarchitecture: Micro-Fusion & Macro-Fusion
- Microarchitecture: Zero (one) idioms & Mov Elimination
- Old Pintool Upgrade with newest pin
interval model¶
ip¶
- IP Forward
- IPV4 && IPV6
- Linux Auto Run : crontab
- Linux Network Command Guide
- ServerLogin
- USTC Network Information Center
ipcc¶
- Hybrid Multithreaded/OpenMP + MPI parallel Programs
- IPCC Preliminary SLIC Analysis
- IPCC Preliminary SLIC Analysis part2 : Run process
- IPCC Preliminary SLIC Analysis part3 : Hot spot analysis
- IPCC Preliminary SLIC Analysis part4 : cluster environment
- IPCC Preliminary SLIC Case1/2/3
- IPCC Preliminary SLIC Optimization 1
- IPCC Preliminary SLIC Optimization 2
- IPCC Preliminary SLIC Optimization 3
- IPCC Preliminary SLIC Optimization 4: EnforceLabelConnectivity
- IPCC Preliminary SLIC Optimization 5: MPI + OpenMP
- IPCC Preliminary SLIC Optimization 6: Non-blocking MPI
- IPCC Preliminary SLIC algorithm
- IPCC Preliminary SLIC test
- Training course - IPCC 5 Optimize common tools
- VtuneOptimize
iptable¶
iptables¶
ipv4¶
ipv6¶
ispc¶
ive¶
izone¶
j4125¶
java¶
jekyll¶
job¶
jpg¶
k-fold cross validation¶
- Pytorch 1 :Basic Components
- Pytorch 2 :more conceptions about training and inference
- Pytorch 3 :Model & Training
- Pytorch 4 :Save & Load & Pretrain
- Pytorch 5 : Distributed Training & Parallelism (Mem Re-allocated)
- Pytorch 6 :Visualization
kill¶
kpop¶
kunpeng 920¶
latex¶
linux¶
live2d¶
llm¶
llvm¶
- Kunpeng
- LLVM Mca : huawei HiSilicon's TSV110 work
- LLVM Mca :with BHive (2019)
- LLVM-MCA: Install&RunTests
- LLVM-MCA: docs
- llvm
- llvm Backend
- llvm Pass
llvm-mca¶
- LLVM-MCA: docs
- Static Code Analysis
- uops.info: Characterizing Latency, Throughput, and Port Usage of Instructions on Intel Microarchitectures (2019)
local-dev¶
localhost¶
logic core¶
loop¶
lscpu¶
lsof¶
macbook¶
make¶
malloc¶
manga¶
map¶
master¶
mca¶
- Kunpeng
- LLVM Mca : huawei HiSilicon's TSV110 work
- LLVM Mca :with BHive (2019)
- LLVM-MCA: Install&RunTests
- LLVM-MCA: docs
memory¶
micro-op fusions¶
minicoda¶
mkdocs¶
moonlight¶
mount¶
mpi¶
mpicc¶
mpiicc¶
mtr¶
nD-FullMesh¶
name¶
nas¶
nasm¶
neon¶
network¶
networkmanager¶
newline¶
nginx¶
npm¶
nsight¶
nvidia¶
nvprof¶
omni¶
oneApi¶
oneapi¶
openmp¶
openmpi¶
openvpn¶
opkg¶
optimization¶
optimize¶
optimizer¶
overleaf¶
page table¶
pam¶
paper¶
- LLVM Mca :with BHive (2019)
- Latent Dirichlet Allocation (2003)
- Micro2023: Utopia
- uops.info: Characterizing Latency, Throughput, and Port Usage of Instructions on Intel Microarchitectures (2019)
parallel¶
password¶
pdf¶
perf¶
piano¶
ping¶
pip¶
png¶
podman¶
port¶
ppt¶
prefetch¶
price¶
process¶
profiling¶
proxy¶
- Clash on LAN/linux/Dockers
- Cloudflare warp proxy
- Introduction to Git Commands
- Network Basic
- SSHForward
ps¶
pull request¶
pyc¶
pyo¶
pypy¶
python¶
- Debug/Profile/Devlop Tools of PTA
- Latent Dirichlet Allocation (2003)
- Pip Package
- Python
- Python Class
- Python Graph & Visualization
- Python MPI
- Python: DataStructure
- PythonRegex
- WebCrawler first try
pytorch¶
qt¶
ram¶
ranking¶
rar¶
rating¶
readelf¶
reboot¶
reduce¶
regex¶
register¶
registers¶
report¶
rl¶
route¶
router¶
rpm¶
safety¶
scons¶
scss¶
security¶
segmentation fault¶
server¶
sh¶
simd¶
skill¶
smanga¶
smb¶
snap¶
snat¶
sniper¶
socket¶
speaking¶
sql¶
sram¶
ssh¶
ssl¶
stack¶
- CSAPP: Machine Programming III: Procedures
- Linux Executable file: Structure & Running
- [C++ Basic] STL Data Structure
stencil¶
step-video¶
stream¶
switch¶
synology¶
systemctl¶
systemd¶
t2v¶
- 250217 Step-Video-T2V Reading & Porting
- 260117 Step-3-VL 10B
- AI Model Memory
- Ideas around T2I2V models
- Ideas around Vision-Language Models (VLMs) / Reasoning Models
- World Model/UFMs/Omni-Modal: AR vs DiT
tar¶
tcp¶
tcpdump¶
terminal¶
thread¶
tick¶
tips¶
tlb¶
tmm¶
tmpfs¶
tools¶
top¶
topology¶
torch¶
torch.dist¶
torch_npu¶
torchrun¶
traceroute¶
uTorrent¶
udp¶
ufw¶
ugreen¶
unix¶
unravel¶
usb¶
user¶
useradd¶
usermod¶
vLLM¶
vLLM-Ascend¶
vLLM-Omni¶
vector¶
vectorization¶
verl¶
video¶
vim¶
virtual machine¶
visualization¶
vllm¶
- Speculative Decoding & eagle3
- Vllm Basic
- vLLM Inference Profiling
- vllm-omni & DiT Inference Accelerate
vlm¶
vpn¶
vscode¶
vscode debug¶
vtune¶
wake¶
warp¶
webdav¶
website¶
- Web Design 1 : Layout Overview
- Web Design 2 : Content Organization & Link Content using go template
- Web Design 3 : Future Features
- Web Design 4 : Customize Markdown Grammar In SSG
wget¶
wifi¶
windows¶
wireguard¶
- ClashX Pro and Wireguard on Macbook In School Net
- Github Access
- OpenWRT on router
- Wireguard
- Wireguard Server 2 Server in OpenWRT