• 1 Post
  • 139 Comments
Joined 3 years ago
cake
Cake day: June 10th, 2023

help-circle
  • Both the tracker and the peers in the swarm could see the percentage that you have, as it is announced for coordination with the swarm. It’s just more data that could be used to profile you, yes. On a very common torrent, it probably won’t matter, since more people will have restarted at similar percentages.

    If you’re worried about it, you can restore the torrent, delete the partial content, force recheck on the torrent and redownload it from zero. It will download the whole torrent again, hurting your ratio, but if the content is 20 GiB and you have uploaded 200 GiB on that torrent, you’ll at least keep a positive ratio on the torrent.

    If your torrents are close to 100% but not quite, because you haven’t downloaded the .nfo or .url file, it won’t matter as much. A lot of people also skip those files, so the download % without those files is very common in the swarm. I still recommend downloading those files, as they barely use up any space, and downloading them helps with the torrent’s health.


  • Those files won’t contain your IP. Your IP is never stored by qbittorrent. Your statistics also never leave qbittorrent. Nobody can request them from you.

    It would technically be possible to profile you, if you had a lot of rare torrents and someone made a list of them. If the same list suddenly appears somewhere else, and the list is sufficiently rare, then it could be assumed that it’s the same person. This can be mitigated by adding the torrents slowly instead of all at once. That way it will look more natural, like someone else slowly downloaded more stuff and ended up with a similar list.

    If you’re curious about the binary data, it’s just the torrent hash tree for integrity checks and identification of the torrent.


  • I’m not actually an expert on the matter, and there is probably someone who knows more.

    There are 2 protocols that qbittorrent uses for data sharing. The oldest one is basic TCP, with the same overhead. The “newer” (it’s already quite old, too) protocol is µTP. It was specifically designed for p2p purposes, specifically torrents. µTP is, among other things, designed to coexist with other traffic without slowing it down. For this purpose, its packet size is variable. On very small packets, the overhead is big, while on larger packets it’s small. I don’t know how different the average overhead is between the 2 protocols, since those variable packets were designed for shared ADSL. Modern broadband probably allows µTP to use bigger packets most of the time.

    From what I can gather, encryption also adds overhead to the traffic.

    This is, however, not the main source of overhead (I don’t even know if that overhead is really counted, I’d have to look more into it). The main source of overhead is actually control traffic (and that’s the reason why you get download even when only uploading). Peers constantly communicate with each other to form the swarm, and they need to request blocks and announce which blocks they have. As far as I understand, this is the main source of the overhead, and it’s independent from the protocol. Your client will announce to other peers what pieces it has, and its capabilities.

    I also don’t know if tracker traffic is included in the count. Your client sends a status to trackers every n minutes, which are specified by the tracker. Those are called announces. Most trackers have announce times between 30 minutes and 1h. If you have 500 torrents, with an average of 5 trackers per torrent and an average announce time of 45 minutes, you’re sending ~3.3k announces per hour. They don’t take much, but they add up.

    DHT and PeX traffic could also be counted, but as with trackers, I don’t know if it is.

    If you start trying to get higher and higher ratios, you’ll notice there’s an asymptotic curve the higher you go. It really seems impossible to get to 30. I once had a client lose its historical up/down data, and without any download, it stopped at a ratio of around 32. This may be the ratio between upload + upload overhead / download overhead (from the PoV of your client).

    Now, while I don’t have any proof and I don’t know how qbittorrent counts overhead, I suspect all those overheads are counted at the same time. So you’re basically stacking them. This is because of the traffic limiter. If you really want to limit traffic to 100mbps, ideally, you’d count every overhead. If you don’t, you could end up with both an unpredictable amount of traffic, and a higher amount of it than 100mbps. If they already implemented a counter that considers all overhead (basically just counting raw packet data), it would make sense that they used it for the global traffic counter.






  • I said I wouldn’t answer, but I just can’t help it. I’m literally laughing at this reply.

    You just got fucking schooled by Chad GPT. Have a blissful day.

    Yeah, I guess that explains why everything is so misinformed and incorrect.

    I’m not even gonna bother answering any of the points, as just 5 minutes of research is enough to disprove them.

    Though, I am gonna share my favourite part of your rant:

    An ASIC doesn’t have to be a single-purpose chip that becomes worthless if you change one training algorithm.

    Yeah, I guess the Application Specific Integrated Circuit doesn’t have to be application specific after all.

    I’ll also share this other gem with you. You wrote it yourself! I just made a few changes to it to make it more accurate.

    “It’s funny how a guy with no engineering experience whatsoever thinks they know more than people designing server systems and computing solutions, who have experience working with datacenters.”



  • You don’t know what you’re talking about. Heat pump circuits don’t generate extra heat? Semiconductors don’t need to stay cool? They don’t need constant temperature? You don’t need accesibility? Dismissing Microsoft’s story because it happened on earth? Quoting Musk’s numbers without questioning if they’re even possible? You think ASICs will fix everything?

    Yes, heat pumps generate heat.

    Yes, semiconductors need to stay cool and have constant temperature.

    Yes, you don’t need to use standard hardware, but then, why spend 100x more on hardware to have 40x less performance? It makes no sense. You can only use appropriate hardware on earth.

    Yes, you need accessibility. Are you gonna deorbit and burn a whole “100-150 kW” satellite just because you needed to change a cable on a patch panel? Microsoft’s story demonstrates how even with free power and cooling, accessibility breaks the deal because it’s that important. And Microsoft “only” had to refloat a container and open a hatch. They didn’t have to burn the whole thing.

    Additionally, burning metals in the atmosphere is incredibly harmful for the ozone layer. So there’s that.

    Even if you built that “gigawatt datacenter” in space, that’s a meaningless term. It only makes some sort of sense on earth, where we know the performance/watt of common hardware. The term is a sensationalized thing, and the correct measure is performance, like flops. A gigawatt datacenter in space, with hardened hardware, will be way slower than a 50 megawatt datacenter on earth.

    You say you could have custom ASICs built for the purpose? While some massive companies have made their own hardware before, look at nvidia. The only reason they’re at the position they’re in is because they have designs for AI accelerators. Only nvidia can design those fast chips, and even if they’re really expensive, most AI datacenters (and even compute intensive non AI datacenters) are full of those. Nobody can build anything, ASIC or not, that beats nvidia.

    ASICs are also a really bad choice for a datacenter, because they’re expensive, have to be custom built for a task, and make your super expensive datacenter useless for everything else. Even changing the training method of the AI will make your ASICs useless.

    And no, you can’t just “build a chip for the environment” and expect to not have to cool it to the same strict levels, or shield it, or have error correction. Semiconductors are not magic, they work in a certain way. And of the litographic processes announced for the so called “terafab”, none of them can be radiation hardened. They’re too small.

    I don’t know why I even bothered to write this, and I don’t think I’ll bother to write the next reply when you inevitably quote more of elon’s delusions at me. You clearly have never been close to a datacenter, and you don’t understand what you’re talking about. You believed musk when he said it totally made sense, and your whole argument makes no sense at any level.


  • That is just a load of bs. Heat pumps do not operate efficiently at high temperatures, and they only add more heat to the circuit. You cannot generate energy from the temperature difference, because then you would slow heat transfer significantly while generating an insignificant amount of power.

    It has already been calculated, and radiators are a really big issue for “space datacenters”. Even with our current cooling technology advancements. Server hardware needs to be kept at a cool temperature (ideally not higher than 50°C), and most importantly, it needs to be constant in spite of very variable hardware load. That is really difficult to manage in space, and an issue that needs to be considered even for our non datacenter satellites.

    But even if cooling was trivial, space datacenters make absolutely no sense.

    For starters, it wouldn’t be possible to put any server in the market up there. Chips are very vulnerable to the radiation in space, which is why space CPUs are stupidly expensive and way slower than earth CPUs. They need to use less efficient litographies, with very special shielding that isn’t commonly available. The sheer compute and memory density of our current servers would be impossible to properly shield, even if we ignore the cost. ECC RAM and ZFS pools for drives would very quickly stop protecting your server from errors, since the bit flips would be too common, and your computations would often be incorrect.

    If you ignore both cooling and space radiation, you still have the issue of accesibility. Datacenters have people always coming and going, because hardware needs maintenance. Sometimes you want to upgrade the server, change the connections, add or remove hardware, and of course, you need to replace the parts that break over time. Hardware becomes old really quickly, especially for the compute intensive workloads that are planned for those “space datacenters”.

    Datacenters are kept fast and competitive through constant upgrading of the parts. Compute intensive servers also churn through parts very quickly, because even at 50-60°C temps, keeping hardware at 100% utilization isn’t the best. This is further exhacerbated by temperature fluctuations, which can prematurely kill the silicon.

    Like someone else mentioned in another comment, Microsoft already tried dumping containers with servers into the ocean, and putting wind turbines on top. That’s free cooling, very stable temperatures, and free energy. Even if it was a success in every measure, Microsoft dropped it because it wasn’t accessible enough, which made it unusable for any real usecase.

    The truth is, datacenters are a market where cost is very important, and solutions are already very competitive. Putting a datacenter in orbit would be stupidly costly, even if you could fix all the issues mentioned. Free electricity does not offset those costs, and that’s without taking into account the cost of the rockets for putting the datacenters in orbit.

    TL;DR: There are so many issues and extra costs, even when ignoring cooling, that make datacenters in space a stupid thing to even consider.

    And before you say I’m not an engineer and I don’t know what I’m talking about, server infrastructure is my field. I’ve worked with datacenters, maintained server stacks and I myself have worked on those cost calculations. I know what I’m talking about.




  • Yeah. This method is really cool, and it has been used for legitimate cases before. Look at the EICAR file:

    X5O!P%@AP[4\PZX54(P^)7CC)7}$EICAR-STANDARD-ANTIVIRUS-TEST-FILE!$H+H*
    

    It is a test file for malware detection software (in fact, it will still get flagged today if you try to send it, download it or save it to your disk). It prints “EICAR-STANDARD-ANTIVIRUS-TEST-FILE!” when run.

    It was purposefully made so all the opcodes appeared as characters you can type on a standard keyboard. That way, anyone can write the file themselves even if they don’t have internet or a way to download it. It’s also easier to get it into test containers/VMs that are isolated from the internet or physical hardware, which is where most malware testing happens.




  • Rar is proprietary, so you can’t freely implement it. Only the decompression code is freeware, and you have to implement it as is. The license includes clauses against reverse engineering and rewriting the code.

    Microsoft could totally pay the fees, but you know how scummy they are.



  • Afaik (and I may be wrong), RAR is a proprietary format created by the guy who made winrar. Because of it, the compression/decompression libraries are licensed, and cannot be freely written. However, in order for the format to catch on, the developer released a freeware (free as in beer, not as in freedom) decompression library. That way, anyone can easily incorporate decompression code in their software, and thus, anyone can decompress a rar.

    However, the code isn’t free to rewrite, so rar decompressors (at least those that don’t make a deal and pay money to winrar) are stuck with the same implementation of the rar decompressor. If that freeware code is slower than what winrar uses, winrar will always be faster.

    Imho that’s a bit anti-competitive, and is what has kept me from ever using the RAR format. Especially now that we have way better formats (both in compression and features), like 7z.