I just started a pbs server to check out how it works. I have 8 vm's at 50GB each, and each has 5 snapshots. Each snapshot appears to be the full 50GB and the dashboard is giving me a dedupe factor of 1. I have the backup retention set to last 14. How long before I should see the deduplication saving space?
Just got this up and running yesterday. Set it for daily backups. It did the first full backup and this morning did the first incremental. Shows only a value of 3.04. TechnoTim had a value like 64. Not sure what he was doing for that but I do think he was doing like hourly backups so that maybe is why. Anyhow just curious what things I could do to possibly increase this value?
On that note, does the garbage collection schedule make any difference? Right now I am doing on PBS pruning for 7 dailys and 2 weeklys (pve job retention is off) and garbage collection every 6hrs. Not sure if this impacts anything but wanted to mention.
I guess what is odd (more with the incremental aspect). I haven't changed a single thing on any of my 10 CTs yet it got an incremental backup and took an hour. Still better than the almost 3hrs it too for a full backup but confused why it was so much if it's just the difference. Since nothing changes shouldn't it be close to nothing?
Hi,
so im currently testing the proxmox backup server für my vms. I was wondering if the deduplication / dirtybit only works with zfs or if it also works on "normal" ext4 formated disks (mounted directory).
Best
I'm using Proxmox at home to host the usual things. One ZFS NVMe pool for all the containers/VMs and one ZFS HDD pool for nightly backups.
I just setup a second PC with TrueNAS for some "just in case" backups of the backups (from PVE and other sources around the house).
Each VM has 1-2 dozen nearly identical backups, so I decided to try ZFS' deduplication feature. I copied over backups for 4 VMs (500G total) and was shocked to see no gains/efficiency from the deduplication, not even 1.01!
Some of those backups are from a VM that was powered off, so the original disk was definitely unchanged. Shouldn't the resulting backups be bit-identical? Does anyone know if Proxmox includes some metadata that might spoil block-level deduplication?
I am in the process of migrating from Veeam to PBS, and there are some differences but it all seems managable for my scenario.
One detail I can't seem to find clarified, that maybe someone in here knows: dedup on PBS, is that across a single VM's backups, across whole datastores, or globally?
Today I played around with some settings in ZFS and Proxmox Backup Server. I use Proxmox Backup Server with a ZFS raidz2 as storage for the Backup space. The ZFS pool was created manually without the proxmox-backup-manager command.
I noticed that the WebUI shows a deduplication factor of (in my case) 94.02 but the zfs pool shows dedupratio 1.00x.
Does Proxmox Backup Server not use ZFS' deduplication functionality? Should I enable deduplication and compression in the ZFS pool? Which compression algorithm should I use?
EDIT: I have tested a bit with ZFS and Proxmox Backup Server for quite a while (both hardware and VMs) and ZFS' deduplication and compression have next to 0 gains. As pointed out by the comments deduplication does not make sense as Proxmox stores backups in binary chunks (mostly of 4MiB) and does the deduplication and most of the compression there. ZFS compression with zstd-12 ended up on a compressratio of 1.02x which is not very significant.
TL;DR It's not necessary to use ZFS compression or deduplication because it's both already handled by Proxmox's way of storing Archives and Block Devices.
I believe they do not use zfs's dedup functionality. It seems they have created their own file format, Proxmox File Archive (.pxar). So they handle deduplication and compression and data transport themselves. ZFS isn't a requirement for datastores. You can create datastores on ext4 for example. They store data at the file level, not block.
So, I would turn off compression because it's probably extraneous for achieves which are already compressed.
PBS only dedups on a per group (container/vm/host) basis. So you will have some data duplicated across groups. So enabling ZFS deduplication may help a little by deduping the whole datastore instead of each individual group. I don't know if it's worth it though, you'll have to test/decide for yourself.
ETA: a word
The PBS does not use ZFS deduplication. PBS dedupe works similar to Windows offline Dedupe. PBS does it on the fly instead of as a scan process like Windows. But with both, files and blocks are broken up into chunks and stored in a chunk store. The chunks are hashed and matching hashes are used to dedupe.
I took over the mgt of 6 different hosts and a PBS. The PVE hosts are not in a cluster so standalone and I have a suspicion that, because of the fact that vmids are reused on multiple hosts, backups are not lining up well and might be overwritten or unavailable for restoring.
So my question is: should I always use unique vmids for my vms to keep backups healthy?
Hi everyone,
my datastore is at about 40% usage after deduplication. When I run a new backup however, it quickly fills the whole datastore and the following backups fail.
After making some space I’ve run an Garbage collection which worked last time since the backups failed a few days ago. Situation released the following night and now the GC only marks the chunks as „pending removals“. On the Proxmox Forum a staff member mentioned, that new backups are protected 24h from GC, so that they are save while still being written.
How do I set this up, so that my datastore doesn’t blow up every time I run the backup?
So I've put Proxmox Backup Server into production in my org as we move over to PVE from Hyper-V. It is doing a great job replacing Retrospect and Unitrends. In my Unitrends and Retrospect systems, I can run a report that tells me the rate of deduplication of the data I'm storing so I can see if we're storing a lot of data that is vastly different at the block level, or if we're storing a lot of data that is not changing much at the block level from day to day.
In Proxmox Backup Server, on each of the datastores I have a "deduplication factor" line and then a number. In my case the number is 177.09. The issue is I'm not able to find any documentation on what this number means or how it's calculated. Did I miss something? I can't seem to figure out what is changing it either. It has gone up and down over the past month with no real pattern I can work out.
Any help is appreciated. I've already poured over the documentation for PBS for hours and can't seem to find anything.
Also i would be cool, if pbs could provide more information. I.e. select a set of vms inside a bigger selection (or all) and then gernate a report of the exclusive size for that selection and also the common costs against the bigger set (i.e: if a block of 1mb from my selected backup is shared with 3 "other" backups, then this would account for a quarter mb). Such a report would be very helpfull in making decisions.
So if that is correct, a deduplication factor of 1 would be the lowest it could go?
Am I correct in that the only benefit of PBS is that it provides a file-level restore? Otherwise, for a single node Proxmox VE, is the built-in backup function sufficient for being able to restore a corrupted VM or CT?
I am planning to use a 2TB external HDD connected via USB and create a backup job within ProxMox VE 8.2.2
Thank you for the advice.
Hey there,
i recently have taken up a new job and got handed a proxmox environment with a PBS.
Now the PBS is saving all its files on a QNAP - and i dont trust that old thing so I am currently working on migrating the files.
Now to the actual question:
The current PBS folder on the QNAP is telling me that its about 3.5 TB (1lxc + 11vms - each with multiple versions because of retention rules) in size but if i look into it and into the folders inside of it (ct, host, vm) all files combined are around 600MB.
Is qnap just stupid? are there some form of hidden files hiding in the pbs folder? Are the backups really just that small?
I recently set up Proxmox Backup Server (PBS) and connected it to an NFS share on my NAS. After that, I linked the PBS to my Proxmox Virtual Environment (PVE) and configured backups.
From what I understand, one of the main benefits of PBS is its ability to perform incremental backups. However, when I check the backups, they all appear to be the same size, even though I made no changes to the container.
Did I miss something during the setup process? How can I ensure that the backups are incremental and not full backups each time?
I have used both veeam and proxnox backup. Pbs is very integrated and works well. Veeam is better on space and has better de duplication from what I can tell. What’s generally recommended to backup proxmox?
Side note if you add a second ssd drive to your server don’t use zfs. It crashes the whole server. I had to format the second drive to ext4 for the added space to work for veeam without crashing (virtual drive placed on the ext4).
I got my Proxmox server up and running and I created a NFS share on my DS920+, then mounted it in the PVE gui. I setup nightly snapshots of the VM and LXC running on it to the NFS share/mount. After reading around, I realized that PVE is not backing up any of it's config files or settings. I found shell script to make backups of all PVE configs and save them to the DS920+ and that is nice and all, but why would PVE only backup guests and not any of the core config?
I searched around and came across Proxmox Backup Server and found that I can run it in a VM on my DS920+. Perfect! This should solve my full backups needs! Wrong! It appears that it only backs up the guests, just like PVE. Who would have thought PBS would offer essentially the same backup functionality PVE? I mean it is called Proxmox Backup Server??
What does PBS offer that PVE doesn't as far as backups go?
Beginner here. Recently I started running a PVE with 6 VM+Containers in a home server machine. It works fine for last 4 months. Now I have a spare Intel NUC with a 2TB drive storage connected, installed with PBS as trial.
-
Tried integrating PVE with PBS and able to run VM backup jobs configured in PVE with PBS as detination. This is good. But currently I am able to achieve the same with a network drive as destination. What is the value addition with PBS?
-
Currently I do backup of PVE host with "zfs send" to network drive. I don't see option to backup PVE host through PBS. Am I missing anything else?
-
There is a TrueNAS VM runs on PVE. This VM hosts most of my home data. I would like to backup this NAS drive to PBS. Is there any TrueNAS + PBS integration possible?
-
What are the other features which is good to try and use? I would like to know the value addition in running a PBS.
running 8 containers and 2 vms. daily backup is around 12gb in size even with zstd compression. I see a lot of redundancy in the content of the containers (the base system from the distro), but this can't be leveraged because each container is backed up individually
is there a way of optimizing the size of the daily backups? was thinking of a hard separation of os and data using two volumes , and only running backup on the data volume. but this would mean that I would have to setup the system from scratch /ansible in case of recovery rather than just reading in the archive file
any hints? thanks
I (will) run my homeserver on PVE with an Intel 12th gen (gbit ethernet). It has about 1.5TB in photo's and videos (pretty stable size) + 30+ docker containers & Home Assistant OS VM.
Currently I use Borg / Vorta to backup to my RPI4 + 2 external drives.
Now, I would like to start using PBS (and dedupe) but I would like to hear some thoughts about the options I have:
Run PBS on RPI4 (8gb, ethernet) using PiPBS?
Run Debian on RPI4 and share drives to PVE. Then PVE to run a VM/LXC with PBS (using the shared drives)
Use my 24/7 windows desktop (i3-8100, 16gb, wifi6 connection) to run PBS in a VM
Other?
I'm not a fan of buying something new if I don't have another purpose then 1 nightly backup for it. PiHole, VPN etc I can already run on my homeserver, so not sure what I would need additional hardware for. But if you see a good reason, then please let me know. Thanks!
Some containers and HAOS I would like to be able to restore within a day, but if the bulk of the data takes a couple of days that's no issue. I would roll out the ethernet cable for my windows desktop for emergency situations.
Thoughts?
Just an FYI, this is a homelab, not a business.
So I have been reading through the docs a bit on PBS's ability to backup to another location. So far I have only been able to find the ability to use remotes and sync jobs. I have seen the NFS setup mentioned but I assume that means directly from my main Proxmox servers, not PBS. I noticed it mentioned in docs that it pulls the data from a remote datastore rather than the main datastore sending it out.
I liked this setup for one reason. I always had issues with NFS randomly losing connection with my NAS. I have messed with it for too long to care at this point. I have a dedicated PBS machine with mirrored SSD's. I would like to make backups of it to my NAS without using NFS mounts and also utilizing PBS's features of dedup and efficiency.
So is it feasible to create a second PBS instance on my NAS (as a VM) with its datastore being a dataset or zvol (NAS uses ZFS) and syncing from the main PBS? Pros to this are that I use PBS's features, PBS manages when things are deleted (not the NAS OS), and I can switch between the two PBS instances to restore a VM. I thought about just getting rid of the current PBS machine and putting it on my NAS but I would lose the dual PBS setup. My NAS is backed up to another onsite NAS and to an Offsite NAS at my parents.
Please let me know if there are other options that I am not aware of. I also do understand keeping a NAS separate from my VM's which is normally what I do aside from apps or systems that benefit from having access to the datasets directly rather than over the network (media server, device backups, server backups, etc. Although maybe a direct link between the services server and the NAS could change that).
I run PBS on a $150 fanless mini-PC backing up 21 VMs and 2 containers to a Synology NAS. Dedup factor 25x, 1.77TB used.
The guide covers setup, installation, verification, and pruning jobs, retention policy gotchas, PBS vs Veeam, and my full homelab setup with screenshots.
Some of you helped me last week when I posted about my 8.5-hour verify times on NFS. Based on your feedback, I bought a Lexar SL300 USB SSD and set it up as a second datastore. Full verify dropped from 8h 35m to 1h 54m. Daily verification with Skip Verified now takes 2 seconds.
https://edywerder.ch/proxmox-backup-server/